01 / WRITEStarting draft
02 / MUTATENew challenger
03 / JUDGEFixed rubric
04 / KEEPBetter or stay
NO MANUFACTURED RESULTS
Your next post starts here.
Choose a topic and a number of rounds. Every score and revision shown here comes from an actual model call.
—
LIVE-API DEMONSTRATION · NOT POSTED
Loads the winner and brief into the form. Choose rounds, then press Start refining.
ACCEPTED REVISIONS
—
ROUNDS COMPLETED
—
RUN API SPEND
—
The refinement curve
Fixed-rubric score, not measured engagement.
Best eligibleChallenger
OriginalLatest round
The iteration trail
No discarded drafts hidden. Open a round to inspect it.
What actually makes a challenger win?
- GPT-4.1 mini creates a draft or rewrites the current best, targeting its weaker dimensions.
- Jev independently rates hook, clarity, audience fit, usefulness, shareability, voice and conversation potential against the same fixed descriptive rubric. Weights are fixed for the chosen goal: shares or replies.
- Code enforces format limits. Jev flags unsupported specific claims and engagement bait. These checks can miss things: review before posting.
- A challenger must be eligible and gain more than one rubric point. A separate blind A/B judgment, with random presentation order and no scores, must also prefer it. Ties keep the incumbent.
- Stop at the round limit or after three non-improving rounds. A higher score is not evidence of higher real-world engagement.