Teaching the new 1T-parameter Inkling model from Thinking Machines the art of the dad joke, and what it actually learned.
I RL-trained Inkling, Thinking Machines' 975B open-weights MoE, to tell dad jokes from my laptop, and the jokes aren't the result that matters. The fine-tune (introducing aimatey/PunTune-0.6) beats its base at every thinking-effort setting on a quarter of the tokens: not by learning new wit, but by learning to pick one answer and stop, while base Inkling (and a few tier-1 frontier models) brainstorms brilliantly and fails to commit.
Measured across Inkling's entire thinking-effort dial under a frozen instrument, the pattern is clean: the dial moves everyone a little; training moved our model a level. Its best work happens at the effort it was trained on, at a fraction of everyone else's token bill.
The field wrote its own headlines along the way: Gemini, perpetually the afterthought, balanced thinking against quality better than any model we measured. GPT-5.6 Terra and Opus landed at the bottom of the board — though sibling GPT-5.6 Sol held lower-mid, above base. And Muse Spark, the new model from a Meta many had already written off, held mid-table and won the esoteric-topics track outright, after we fired it as a judge for costing too much.
Why dad jokes? I wanted to learn Inkling's RL stack in public, and a silly, non-verifiable task is a better classroom than cybersecurity: it keeps the work fun, and it's one of the few domains where I could credibly serve as the calibrated judge myself.
Short version, before the charts get dense: the jokes came out fine. The finding was the fun part. Making a trillion-parameter model funnier didn't make it wittier; it made it decisive, and it did that at a fraction of the cost. Knowing which of those two a training run can actually buy, and which it can't, is the difference between aiming a model and hoping at one. That's the part a product team can use.
Every training run pinned Inkling's thinking-effort at 0.9; the dial never moved during RL. Afterward we measured all three models at five effort settings, each point 150 fresh jokes scored by nine blind panel placements on the frozen instrument:
Effort buys a little everywhere: +0.04 to +0.06 across the whole dial, for base and fine-tunes alike. The weight of the result is what training bought instead, and it shows best on the chart Inkling's own launch argued for, the full cost curve:
Why the level moved and the slope didn't is in the discovery section.
Before any numbers: rate three jokes blind, 1 to 5, on the question everything here optimizes: would a real dad be proud to inflict this? One is a classic, one is untrained Inkling, one is the fine-tune.
The loop is pure tinker-cookbook: subclass Env, emit reward scalars, and GRPO trains a rank-32 LoRA on Inkling's 41B active params; the laptop just sends prompts and rewards, at about half a cent a rollout. The infrastructure was never the problem. The problem is that "funny" has no verifier: every reward comes from an LLM judge you can fool, so the project became a study in measuring a subjective objective without getting Goodharted. I got Goodharted repeatedly: by the reward, then the benchmark, then the judges, then the token budgets. Catching them was the actual work; each one is a section below.
Goodhart's law is the water you're swimming in the moment you train anything: when the real goal is too hard to measure directly, you optimize a proxy, and the model finds every loophole in the proxy while the real goal quietly rots. If you've shipped product, you've already met it. You wanted satisfaction; the only number on hand was time on site or click-through; so the team dutifully lifted that number with louder CTAs, thinner pages, and an ever more wandering UI, and the customer left less satisfied than before. Training dad jokes is the same move with the masks off: I wanted funny, the reward could only see a rubric, and the model learned to farm the rubric. Naming the proxy is easy. Noticing it has quietly detached from the goal is the actual skill, and it's the one every product team that wades into training has to build.
Fifty prompts, five shapes, from single words to entire news articles pasted in as the topic. Click through what the models were actually asked, with sample results measured on the final instrument:
A 20-item, 100-point dad-joke rubric scored by an LLM judge. The pre-training audit set the tone: AI slop wrapped around a pun scored 86. Anchored and trained, the held-out score climbed 11 points, all of it floor-raising while the craft items fell; a third of the rubric was safety points every clean joke gets for free.
✗ defeated by the points we left lying aroundDouble-weighted pun quality across 6,400 rollouts; the pun item moved +0.02. GRPO stalled exactly where the judge's resolution ran out.
✗ defeated by judge resolutionA cheap probe showed a judge's absolute scores swing wildly between sessions while its orderings stay stable. One within-group ranking signal later, pun quality finally moved, the first real movement of the project.
✗ but a surface-similarity diversity tax punished the genre's own "Why did the X…?" formatRebuilt around validated signals only (canon-anchored holistic ranking, brevity priced in code, a single-punchline gate) under the first exam-selected judge (chosen by correlation with my blind ratings, not by reputation).
✓ first human-rated 4 ever · blind mean 2.56 vs the classics' 3.3Replaced the judge's "feels familiar" flag with a fame index over 10,561 real jokes from 5 corpora: famous jokes gutted, obscure parallel invention rides free. Cleanest training dynamics of the project.
✓ the flagship checkpoint; mid-table against the frontier belowThis round was my objection. I pushed back on my agent's minimal reward: "20 weighted criteria should beat 3 signals; a big model can handle it." We pre-registered predictions and ran it. My falsifier fired: 3.22 blind, above the agent's predicted ceiling. The trajectory showed why: every criterion all eight siblings satisfy saturates to zero within-group variance, and the big rubric quietly reduced itself to the comparative signal. Mine worked; its mechanism explained why. Best $75 of the project.
✓ statistically tied with R5 under both instruments; different roads, same modelAfter four iterations the rubric read 20 points above baseline. Then the blind test: the engineered reward was negatively correlated with human taste, and word count alone beat the entire 20-item judge apparatus.
This is a training project, so here is the training. Six GRPO runs, 100 steps each, LoRA rank-32 on the full 975B MoE, 64 episodes per step. Thin line = raw per-step mean reward; bold = 5-step moving average. Reward scales are not comparable across panels: each round optimizes a different reward function (that is the whole story of the project), but the shapes are: every run learns; what changed round to round is what "up" meant.
Human blind ratings across the arc: 1.33 → 1.33 → 2.00 → 2.33 → 2.44 → 2.56, and then the first floor-3 sheet in project history: fresh R5 at 3.33, R6 at 3.22, against the classics' 3.50 (one sitting, n=9 per arm, but not a single joke below "solid groaner").
Building the benchmark turned out to be harder, and more instructive, than training the model. Every fix below was forced by a specific number, most of them caught by me pressing at the seams (usually from the billing dashboard, of all places). The blow-by-blow, including the judge exams and the firing of our most expensive judge, is in the technical appendix.
Two clean-instrument tournaments (1,800 fresh jokes, exam-seated judges, compute parity) settled the field questions: Gemini's crown is brains, not budget (rationed to 1/12th its default thinking, it improved), and the GPT/Opus basement is a style the judges measure worst, not a missing capability. And the selection pattern isn't unique to Inkling: Sol produced some of the tournament's best raw wordplay ("to keep its naval near its navel"; "defying gratuity") and buried it inside podcast-and-breakup templates. The full boards and style receipts live in the technical appendix; what matters here is what the same instrument said about our model.
Under the frozen ladder, our checkpoints dropped: PunTune 0.6 lands #7 of 13 (0.456) and PunTune 0.5 #8 (0.421) — a lift over base Inkling (0.374) of just +0.082 and +0.047, down from ~+0.29 on the old field-relative instrument; a chunk of our earlier mid-pack standing was the old judges' affection for our minimal style. Deeper still, base Inkling's best jokes are excellent; its problem was rambling into truncation, not wit. The honest restatement of what RL bought: reliability and economy, not humor creation. Exhibit A: the fine-tune's most celebrated pun, and base Inkling inventing the same one on its own:
I only play Fortnite every two weeks — it's a fortnight.
I tried to build a base in Fortnite, but it collapsed after exactly two weeks—turns out my fort only lasted a fort-night.
Same pun, found twice: the joke was in the priors; training taught the model to say it in
ten words and stop. And here is where the calibrated human earns his keep again, because
I think the judges' favorite is barely a joke at all. Fortnite is literally named
after the word fortnight; the original mode was holding out for fourteen days. The celebrated
pun is the game's own etymology read back out loud. It surprises no one. It makes you
blank-stare, not groan. A Fortnite joke that earns the groan looks more like: You know
what they call Fortnite in Paris? A battle royale with cheese.
(Mine. The panel never got
to rank it.) The instrument measured selection and economy correctly; whether the selected
thing was ever funny remains the human's call. Which is the standing dispute of the whole
project: my blind ratings put our fresh jokes near the classics (3.2 to 3.3 vs 3.5) while
the neutral panel ranks them mid-table. When your instruments disagree, you don't pick a
winner. You collect more humans:
What's the difference between OpenAI and Anthropic? One's open, the other's Claude.
^ the largest judge-human gap we ever measured ("it only hits on shape"). Every showcase on this page carries its provenance because of this joke.
Strip everything else away and the comparison this project exists for is three models and one dial: base Inkling, and two fine-tunes trained at a fixed thinking-effort of 0.9, all swept across Inkling's effort range on fresh jokes. The frontier models are context; this is the experiment. Pre-registered predictions: extra deliberation should help the base model, and the brevity-trained fine-tunes probably had the dial trained flat. Both were wrong, in the most interesting direction available:
(The curves are at the top of the page; this section is why they can be trusted.)
The broader hypothesis this earns, and the one worth testing beyond dad jokes: outcome-only RL may teach a model to convert its existing deliberation into a committed final answer, without ever supervising the reasoning itself. That, not the joke quality, is the finding I'd take to a research team. And the model result is only half the deliverable: the reusable contribution is the instrument itself: exam-selected neutral judges, compute-matched on both sides of the bench, permanent reference ladders with documented censoring, able to measure any future checkpoint for about $5 without moving the rest of the field.
None of the above is trustworthy on a leaderboard that moves when the competition changes. So the final instrument freezes a ladder of reference jokes per prompt (frontier material + corpus classics; our models ineligible), and every new joke is placed against those frozen rungs in nine blind, order-randomized panel rankings. Scores are permanent units: adding a model, or a new training checkpoint, never moves anyone else's number. The instrument is certified rather than assumed (reliability projects to ~0.74 at nine placements; known floor-censoring documented), and the ruler replicated across two fresh samples: PunTune 0.6 scored 0.461 during certification and 0.456 on fully fresh jokes a day later. Full construction, certification numbers, and limitations are in the technical appendix.
For context (and it is context, not the headline), the full field measured by the same instrument, including muse-spark-1.1, the judge we fired for cost, entered as a contestant and measured fairly by its replacements:
The board's honest asterisk, from the pre-publication audit: under an equal
reasoning ceiling, realized uptake varies ~100×, and it tracks board position: the
top five think 700-1,300 tokens per joke; Opus and Terra think approximately zero
(Opus's entire deliberation on one joke: Let me come up with an original dad joke about
garden hoses.
). The zero-thinkers also write the two-beat minimal style our
off-distribution validation flagged as the judges' blind spot; Terra's bottom ten contains
genuinely competent economy jokes. So: the basement's existence is robust; its
internal ordering is low-confidence. Equal permission is not equal compute, and no
API parameter can make a model want to think, a claim verified in triplicate: through
OpenRouter at every effort setting, then natively against OpenAI (effort-high buys Terra 9
thinking tokens on a joke) and Anthropic, where it turns out Opus 4.8's thinking is
adaptive-only: the model's own classifier decides what deserves deliberation, and
there is no override at any price. Which frames this project's headline neatly: Opus
keeps the effort dial to itself; Inkling hands it to you, and our fine-tune is a story
about what training can do with a dial the user is allowed to touch.
A neutral panel says PunTune, our fine-tune, is mid-table. The human who trained the judges says its jokes sit near the classics. That dispute has exactly one honest resolution: more humans. The arena serves blind rounds (banked from the benchmark, or generated live on your topic by six models including the fine-tune sampling from Tinker), and every ranking feeds a public Bradley–Terry leaderboard.
Written in the spirit of a good bug report, from someone who wants the platform to win.
Arena volume: the me-versus-panel dispute above is the live experiment; every blind vote is data.
Span-level rewards (R7): feasibility verified; the granularity experiment my own objection earned.
Audience conditioning: same topic, different room (your crush · next to grandma · a bee-decline news story). No reward we built measures the room. Highest ceiling on the board.
Best-of-N Sol: if Sol's problem is selection, a cheap ranker over 3 samples should move it half the board. Falsifiable, ~$15.
Visual dad jokes: Inkling is natively multimodal, and this benchmark never touched it. Next track: single-panel cartoons where the model writes the dad-joke caption. Same instrument, new modality.
Audio and delivery: tone is half of a groaner. A spoken study needs real thought about what a measured "dad cadence" even is, so it gets its own arc rather than a bullet here.
Is it Inkling or the recipe? Same reward on other open bases: polishing priors, or installing skill? (The fortnight twins above suggest an answer worth testing.)
Thirty hours in, mid-discussion about whether to add a DPO stage after GRPO. Verbatim from the session log:
I'm Adam. Twenty years across product, growth, and applied experimentation, learning research fundamentals the way I know how: by shipping. I've been training models this way for a couple of years, always informally: find something models are bad at that takes taste (humor, likeness, timing, voice), pick a version of it I actually find fun, then build, fail, and build again until it works. (An earlier cycle produced a goat-scream API. I won't explain it.)
This isn't my first end-to-end RL project; it's the first one I documented start to finish, specifically to share how I learn in a way others can emulate. I picked a deliberately silly topic on purpose: it keeps the learning fun, humor has no verifier so the measurement problem is real, and it's a domain where I could credibly make the quality calls myself.
That calibration mattered; most of the instrument overhauls above started with me refusing to accept a result that didn't pass the sniff test. The 48-entry research diary records every design decision, failure, retraction, and correction along the way, the ones I caused as well as the ones my AI pair caused and confessed to. The methodology got more rigorous because while I may not have a PhD in Machine Learning, I have an honorary one in experimentation and figuring shit out.
So: eval design with examined judges, provenance-verified blind tests, pre-registered experiments, honest caveats on every claim, and each failure became an applied insight. If you're a PM or an engineer wondering whether this world is reachable, publishing all of it is the point.