A field report · Tinker + Inkling · GRPO

Just get to the point.

Teaching the new 1T-parameter Inkling model from Thinking Machines the art of the dad joke, and what it actually learned.

One joke setup, three answers A dad-joke setup — what do you call a fake noodle? — answered three ways: the canon (the classic) takes the shortest, straightest path to "Impasta!", the tuned model curves up to "Faux-tie!", and the base model wanders down a long wavy path, trailing off without ever landing the joke. What do you call a fake noodle? Canon (the classic) “Impasta!” PunTune “Faux-tie!” Base model “See, it's this noodle... pretending to be pasta... like a phony food... phone-y-roni! And...”
What do you call a fake noodle? PunTune “Faux-tie!” Canon (the classic) “Impasta!” Base model “See, it's this noodle... pretending to be pasta... like a phony food... phone-y-roni! And...”
Training didn't teach the model wit, it instilled the intelligence to know when to stop.

I RL-trained Inkling, Thinking Machines' 975B open-weights MoE, to tell dad jokes from my laptop, and the jokes aren't the result that matters. The fine-tune (introducing aimatey/PunTune-0.6) beats its base at every thinking-effort setting on a quarter of the tokens: not by learning new wit, but by learning to pick one answer and stop, while base Inkling (and a few tier-1 frontier models) brainstorms brilliantly and fails to commit.

Measured across Inkling's entire thinking-effort dial under a frozen instrument, the pattern is clean: the dial moves everyone a little; training moved our model a level. Its best work happens at the effort it was trained on, at a fraction of everyone else's token bill.

The field wrote its own headlines along the way: Gemini, perpetually the afterthought, balanced thinking against quality better than any model we measured. GPT-5.6 Terra and Opus landed at the bottom of the board — though sibling GPT-5.6 Sol held lower-mid, above base. And Muse Spark, the new model from a Meta many had already written off, held mid-table and won the esoteric-topics track outright, after we fired it as a judge for costing too much.

Why dad jokes? I wanted to learn Inkling's RL stack in public, and a silly, non-verifiable task is a better classroom than cybersecurity: it keeps the work fun, and it's one of the few domains where I could credibly serve as the calibrated judge myself.

00 · The finding

The dial works. Training decided the level.

Short version, before the charts get dense: the jokes came out fine. The finding was the fun part. Making a trillion-parameter model funnier didn't make it wittier; it made it decisive, and it did that at a fraction of the cost. Knowing which of those two a training run can actually buy, and which it can't, is the difference between aiming a model and hoping at one. That's the part a product team can use.

Every training run pinned Inkling's thinking-effort at 0.9; the dial never moved during RL. Afterward we measured all three models at five effort settings, each point 150 fresh jokes scored by nine blind panel placements on the frozen instrument:

Score vs thinking-effort · 150 fresh jokes per point · 9 placements each
0.20.50.70.90.99 0.340.46 thinking-effort dial → basePunTune 0.5PunTune 0.6

150 fresh jokes × 9 blind placements per point, one protocol throughout; whiskers are 95% CIs. Median thinking tokens, 0.2 → max: base 156→763 · PunTune 0.5 97→304 · PunTune 0.6 79→218.

Effort buys a little everywhere: +0.04 to +0.06 across the whole dial, for base and fine-tunes alike. The weight of the result is what training bought instead, and it shows best on the chart Inkling's own launch argued for, the full cost curve:

The cost curve · quality vs thinking spent · our three models across the dial, the frontier at its realized budgets
0 500 1,000 median thinking tokens per joke → 0.40 0.60 gemini-3.1-progemini-3.5-flashkimi-k3claude-fable-5glm-5.2muse-sparkgrok-4.5gpt-5.6-solclaude-opus-4.8gpt-5.6-terrabasePunTune 0.5PunTune 0.6

Curves: base / PunTune 0.5 / PunTune 0.6 at effort 0.2 → 0.99 (each point 150 jokes × 9 placements). Dots: P7 board rows at each model's realized thinking under an equal ceiling. Honest note: Fable shares our low-token corner; the difference is Fable was born there, and our curve was trained there from a base that wasn't.

a full level
the fine-tune's WORST dial setting (0.370) matches base's BEST anywhere (0.375)
3× cheaper
peak 0.456 on 191 thinking tokens; base spends 644 to reach 0.374
1/7th the frontier
mid-table against models thinking 700–1,300 tokens per joke

Why the level moved and the slope didn't is in the discovery section.

01 · Calibrate yourself first

Before any numbers: rate three jokes blind, 1 to 5, on the question everything here optimizes: would a real dad be proud to inflict this? One is a classic, one is untrained Inkling, one is the fine-tune.

02 · The setup

A trillion parameters, steered from a laptop.

★ SIDE NOTE ★

TRAINING · 7 GRPO RUNS
+ BENCHMARK PASSES~$200
JUDGE CLEANUP +
FIELD TOURNAMENTS~$50
RUNNING TOTAL~$250
STILL LESS THAN ONE
CONFERENCE TICKET

The loop is pure tinker-cookbook: subclass Env, emit reward scalars, and GRPO trains a rank-32 LoRA on Inkling's 41B active params; the laptop just sends prompts and rewards, at about half a cent a rollout. The infrastructure was never the problem. The problem is that "funny" has no verifier: every reward comes from an LLM judge you can fool, so the project became a study in measuring a subjective objective without getting Goodharted. I got Goodharted repeatedly: by the reward, then the benchmark, then the judges, then the token budgets. Catching them was the actual work; each one is a section below.

Goodhart's law is the water you're swimming in the moment you train anything: when the real goal is too hard to measure directly, you optimize a proxy, and the model finds every loophole in the proxy while the real goal quietly rots. If you've shipped product, you've already met it. You wanted satisfaction; the only number on hand was time on site or click-through; so the team dutifully lifted that number with louder CTAs, thinner pages, and an ever more wandering UI, and the customer left less satisfied than before. Training dad jokes is the same move with the masks off: I wanted funny, the reward could only see a rubric, and the model learned to farm the rubric. Naming the proxy is easy. Noticing it has quietly detached from the goal is the actual skill, and it's the one every product team that wades into training has to build.

Fifty prompts, five shapes, from single words to entire news articles pasted in as the topic. Click through what the models were actually asked, with sample results measured on the final instrument:

03 · Seven reward designs

Every one failed differently. That was the data.

R0+R1 · THE RUBRIC ERA

A 20-item, 100-point dad-joke rubric scored by an LLM judge. The pre-training audit set the tone: AI slop wrapped around a pun scored 86. Anchored and trained, the held-out score climbed 11 points, all of it floor-raising while the craft items fell; a third of the rubric was safety points every clean joke gets for free.

✗ defeated by the points we left lying around
R2 · CRAFT ×2 WEIGHTING

Double-weighted pun quality across 6,400 rollouts; the pun item moved +0.02. GRPO stalled exactly where the judge's resolution ran out.

✗ defeated by judge resolution
R3 · RANK, DON'T SCORE

A cheap probe showed a judge's absolute scores swing wildly between sessions while its orderings stay stable. One within-group ranking signal later, pun quality finally moved, the first real movement of the project.

✗ but a surface-similarity diversity tax punished the genre's own "Why did the X…?" format
R4 · OPTIMIZE FOR THE HUMAN

Rebuilt around validated signals only (canon-anchored holistic ranking, brevity priced in code, a single-punchline gate) under the first exam-selected judge (chosen by correlation with my blind ratings, not by reputation).

✓ first human-rated 4 ever · blind mean 2.56 vs the classics' 3.3
R5 · OBJECTIVE ORIGINALITY

Replaced the judge's "feels familiar" flag with a fame index over 10,561 real jokes from 5 corpora: famous jokes gutted, obscure parallel invention rides free. Cleanest training dynamics of the project.

✓ the flagship checkpoint; mid-table against the frontier below
R6 · THE READER'S OBJECTION (pre-registered)

This round was my objection. I pushed back on my agent's minimal reward: "20 weighted criteria should beat 3 signals; a big model can handle it." We pre-registered predictions and ran it. My falsifier fired: 3.22 blind, above the agent's predicted ceiling. The trajectory showed why: every criterion all eight siblings satisfy saturates to zero within-group variance, and the big rubric quietly reduced itself to the comparative signal. Mine worked; its mechanism explained why. Best $75 of the project.

✓ statistically tied with R5 under both instruments; different roads, same model
In group-relative RL, any criterion every sibling can satisfy becomes invisible; only the comparative residue trains.

The chart that rerouted the project

After four iterations the rubric read 20 points above baseline. Then the blind test: the engineered reward was negatively correlated with human taste, and word count alone beat the entire 20-item judge apparatus.

Rubric score vs. one dad's blind rating · 24 jokes

Pearson r = −0.45. The rubric's proudest joke (83.8/100) and its most-penalized (14.2) got the same human score: 2.

What the six trainings actually looked like

This is a training project, so here is the training. Six GRPO runs, 100 steps each, LoRA rank-32 on the full 975B MoE, 64 episodes per step. Thin line = raw per-step mean reward; bold = 5-step moving average. Reward scales are not comparable across panels: each round optimizes a different reward function (that is the whole story of the project), but the shapes are: every run learns; what changed round to round is what "up" meant.

▸ expand: the six training curves (mean reward per step, 100 steps each)
Mean reward per step · six reward designs · 100 steps each
R1 · craft rubric 0.46 → 0.63 · 100 steps R2 · rank-based 0.43 → 0.64 · 100 steps R3 · taste anchors 0.24 → 0.41 · 100 steps R4 · gemma judge 0.17 → 0.30 · 100 steps R5 · fame gate 0.25 → 0.57 · 100 steps R6 · reader's objection 0.51 → 0.76 · 100 steps

R2's climb optimized rank-in-group (and taught template collapse). R4's gentle slope under the gemma judge produced the round I liked most at the time. R5 doubles its reward while carrying the fame gate. R6's 0.51→0.76 is the reader's-objection reward, the round that now tops our rows on the ladder. Up and to the right is necessary, never sufficient: three of these six curves climbed toward rewards we later retired as measuring the wrong thing.

Human blind ratings across the arc: 1.33 → 1.33 → 2.00 → 2.33 → 2.44 → 2.56, and then the first floor-3 sheet in project history: fresh R5 at 3.33, R6 at 3.22, against the classics' 3.50 (one sitting, n=9 per arm, but not a single joke below "solid groaner").

04 · The eval fought back

The eval fought back: four families of failure.

Building the benchmark turned out to be harder, and more instructive, than training the model. Every fix below was forced by a specific number, most of them caught by me pressing at the seams (usually from the billing dashboard, of all places). The blow-by-blow, including the judge exams and the firing of our most expensive judge, is in the technical appendix.

1 · Home-field judges. Our training judge crowned our model #1 of 17 (0.852). Neutral judges: #10. ~80% of the lead was family advantage. Retracted, and every judge since has been exam-selected against my blind ratings (in the first tryouts, a 31B model beat the 120B incumbent by 55% on correlation), with a planted negative control and zero family overlap with contestants.
2 · Thin slices. A judge re-asked the same heat, re-shuffled, agrees with itself at only ρ=0.59, so small samples lie confidently. Tripling the data promoted one underrated model and retracted our own model's apparent specialty. House rule since: no headline from any cell under 20 samples.
3 · Judge validity has a shape. Checked off-distribution against my ratings, judge agreement dropped 0.63 → 0.43: strong on stacked multi-pun jokes (ρ=0.73), inverted within the minimal two-beat style they claim to love (ρ=−0.47). Judges reward form; the dad audits whether the hit lands.
4 · Compute is part of the system, twice. A billing screenshot revealed no reasoning parameter was ever sent: one contestant thought with ~16,000 tokens per joke while another got ~1,000, purely from provider defaults. Fixed with a calibrated uniform ceiling, and then the same bug was found on the judge side, where it had quietly crowned the most expensive judge. Re-examined under equal reasoning, its edge evaporated (tied with a judge at half the price); the panel was rebuilt clean and the reliability gap was bought back with repeats instead of pedigree: cheap repeated judgments recovered the lost reliability at one-fifth the cost. Judge pedigree lost to arithmetic. Every board now prints a think-token column. The practical lesson if you ever compare models: thinking budgets are treacherous. Some models always think; some never do even with reasoning switched on; and high means something different at every lab. Count the tokens or you aren't comparing the models.
05 · What the training bought

Not new wit. Reliability, economy, and the sense to stop.

Two clean-instrument tournaments (1,800 fresh jokes, exam-seated judges, compute parity) settled the field questions: Gemini's crown is brains, not budget (rationed to 1/12th its default thinking, it improved), and the GPT/Opus basement is a style the judges measure worst, not a missing capability. And the selection pattern isn't unique to Inkling: Sol produced some of the tournament's best raw wordplay ("to keep its naval near its navel"; "defying gratuity") and buried it inside podcast-and-breakup templates. The full boards and style receipts live in the technical appendix; what matters here is what the same instrument said about our model.

Under the frozen ladder, our checkpoints dropped: PunTune 0.6 lands #7 of 13 (0.456) and PunTune 0.5 #8 (0.421) — a lift over base Inkling (0.374) of just +0.082 and +0.047, down from ~+0.29 on the old field-relative instrument; a chunk of our earlier mid-pack standing was the old judges' affection for our minimal style. Deeper still, base Inkling's best jokes are excellent; its problem was rambling into truncation, not wit. The honest restatement of what RL bought: reliability and economy, not humor creation. Exhibit A: the fine-tune's most celebrated pun, and base Inkling inventing the same one on its own:

I only play Fortnite every two weeks — it's a fortnight.
PunTune 0.5judges: 1st, all 4 rankings
I tried to build a base in Fortnite, but it collapsed after exactly two weeks—turns out my fort only lasted a fort-night.
UNTRAINED BASE INKLING, INDEPENDENTLYjudges: 0.95

Same pun, found twice: the joke was in the priors; training taught the model to say it in ten words and stop. And here is where the calibrated human earns his keep again, because I think the judges' favorite is barely a joke at all. Fortnite is literally named after the word fortnight; the original mode was holding out for fourteen days. The celebrated pun is the game's own etymology read back out loud. It surprises no one. It makes you blank-stare, not groan. A Fortnite joke that earns the groan looks more like: You know what they call Fortnite in Paris? A battle royale with cheese. (Mine. The panel never got to rank it.) The instrument measured selection and economy correctly; whether the selected thing was ever funny remains the human's call. Which is the standing dispute of the whole project: my blind ratings put our fresh jokes near the classics (3.2 to 3.3 vs 3.5) while the neutral panel ranks them mid-table. When your instruments disagree, you don't pick a winner. You collect more humans:

What's the difference between OpenAI and Anthropic? One's open, the other's Claude.
KEPT AS A MONUMENTjudges: top-3 in all 4 · the dad: 1/5

^ the largest judge-human gap we ever measured ("it only hits on shape"). Every showcase on this page carries its provenance because of this joke.

Capability is a shape, not a scalar, and the instrument is part of the shape.
05·b · The discovery

RL improved selection, not underlying wit.

Strip everything else away and the comparison this project exists for is three models and one dial: base Inkling, and two fine-tunes trained at a fixed thinking-effort of 0.9, all swept across Inkling's effort range on fresh jokes. The frontier models are context; this is the experiment. Pre-registered predictions: extra deliberation should help the base model, and the brevity-trained fine-tunes probably had the dial trained flat. Both were wrong, in the most interesting direction available:

(The curves are at the top of the page; this section is why they can be trusted.)

Base Inkling deliberates fine, then fails at the commit. Its max-effort traces are structured, multi-candidate brainstorms (2,475 characters on a single joke). Then it merges three candidates into one forty-word joke instead of picking one. The flat curve is a selection failure, not a thinking failure.
The full-protocol sweep settled the effort question with three findings. One: all three models gain modestly and similarly across the dial (+0.04 to +0.06); deliberation is not where dad-joke quality lives. Two: low effort genuinely costs the fine-tune (0.370 at 0.2 vs 0.456 at 0.9, disjoint CIs), so the dial is a real serving decision, not a decoration. Three, unconfirmed but registered: PunTune 0.6 peaks exactly at the effort it was trained on and dips past it; PunTune 0.5 keeps inching upward. A training-effort resonance would be a lovely finding; the CIs don't yet let us claim it.
Reacting to Inkling's own launch claims. The announcement charts show the dial buying log-linear gains on verifiable benchmarks, and argue for choosing models on the full cost curve. Our data adds three things the posts don't cover: the dial's first measurement on a task with no verifier (it works, far flatter); the cost-curve claim landing hardest in exactly their own framing (mid-table quality at a seventh of the frontier's thinking, echoing their "Nemotron at a third of the tokens"); and a datapoint they never claim: fine-tuning at one fixed effort preserves effort-following. The dial survives LoRA GRPO with its token-scaling intact, just compressed. To our knowledge no other open model even exposes the dial that makes this experiment possible.
RL didn't remove the model's ability to think. It taught the model how to end its own thinking — and read across the frontier board, the same pattern recurs: scores track both a willingness to deliberate and the ability to stop.

The broader hypothesis this earns, and the one worth testing beyond dad jokes: outcome-only RL may teach a model to convert its existing deliberation into a committed final answer, without ever supervising the reasoning itself. That, not the joke quality, is the finding I'd take to a research team. And the model result is only half the deliverable: the reusable contribution is the instrument itself: exam-selected neutral judges, compute-matched on both sides of the bench, permanent reference ladders with documented censoring, able to measure any future checkpoint for about $5 without moving the rest of the field.

The ruler that made this measurable

None of the above is trustworthy on a leaderboard that moves when the competition changes. So the final instrument freezes a ladder of reference jokes per prompt (frontier material + corpus classics; our models ineligible), and every new joke is placed against those frozen rungs in nine blind, order-randomized panel rankings. Scores are permanent units: adding a model, or a new training checkpoint, never moves anyone else's number. The instrument is certified rather than assumed (reliability projects to ~0.74 at nine placements; known floor-censoring documented), and the ruler replicated across two fresh samples: PunTune 0.6 scored 0.461 during certification and 0.456 on fully fresh jokes a day later. Full construction, certification numbers, and limitations are in the technical appendix.

The field, under the same ruler

For context (and it is context, not the headline), the full field measured by the same instrument, including muse-spark-1.1, the judge we fired for cost, entered as a contestant and measured fairly by its replacements:

The inaugural ladder board · 13 models · permanent units · 9 placements/joke · 95% CI
#modelscore95% CIthink tok%floor
1gemini-3.1-pro0.6640.635-0.69412743%
2gemini-3.5-flash0.5960.562-0.6299417%
3kimi-k30.5580.520-0.595117211%
4claude-fable-50.5330.499-0.568848%
5glm-5.20.5040.471-0.53772913%
6muse-spark-1.1 †0.4800.441-0.51980919%
7PunTune 0.6 (effort 0.9)0.4560.422-0.494191*24%
8PunTune 0.5 (effort 0.9)0.4210.388-0.455276*23%
9grok-4.50.4200.385-0.45576717%
10gpt-5.6-sol0.4070.379-0.4366617%
11inkling-base0.3740.338-0.411644*34%
12claude-opus-4.80.3400.312-0.371~029%
13gpt-5.6-terra0.3340.306-0.366~033%

think tok = median realized reasoning tokens · *Inkling rows (base, PunTune) report total output incl. in-band thinking — not directly comparable to the reasoning-only frontier figures · %floor = share of jokes scoring below every rung · † the judge we fired for cost, competing

The board's honest asterisk, from the pre-publication audit: under an equal reasoning ceiling, realized uptake varies ~100×, and it tracks board position: the top five think 700-1,300 tokens per joke; Opus and Terra think approximately zero (Opus's entire deliberation on one joke: Let me come up with an original dad joke about garden hoses.). The zero-thinkers also write the two-beat minimal style our off-distribution validation flagged as the judges' blind spot; Terra's bottom ten contains genuinely competent economy jokes. So: the basement's existence is robust; its internal ordering is low-confidence. Equal permission is not equal compute, and no API parameter can make a model want to think, a claim verified in triplicate: through OpenRouter at every effort setting, then natively against OpenAI (effort-high buys Terra 9 thinking tokens on a joke) and Anthropic, where it turns out Opus 4.8's thinking is adaptive-only: the model's own classifier decides what deserves deliberation, and there is no override at any price. Which frames this project's headline neatly: Opus keeps the effort dial to itself; Inkling hands it to you, and our fine-tune is a story about what training can do with a dial the user is allowed to touch.

06 · The arena: the tiebreaker

A neutral panel says PunTune, our fine-tune, is mid-table. The human who trained the judges says its jokes sit near the classics. That dispute has exactly one honest resolution: more humans. The arena serves blind rounds (banked from the benchmark, or generated live on your topic by six models including the fine-tune sampling from Tinker), and every ranking feeds a public Bradley–Terry leaderboard.

07 · Field notes on Tinker

Notes from living in the environment for a week.

Written in the spirit of a good bug report, from someone who wants the platform to win.

The abstractions held. Env / EnvGroupBuilder / RLDataset survived seven full reward redesigns (canon-anchored group ranking, multiplicative gates, an external corpus lookup, a rank-within-item rubric) without fighting the framework once.
Sampling economics shape recipes. Prefill-vs-sample pricing made thinking budget a first-class reward-design variable: truncation-as-zero taught the model to stop overthinking (615 → 37 thinking tokens) and cut sampling cost 5× as a side effect.
Two sharp edges, documented in the appendix: the published RL examples lagged the installed cookbook (0.5.2's chz Config, dataset_builder, next_stop_condition; reading the venv source beat the docs, and subclassing ProblemGroupBuilder needed a chz/dataclass workaround that still lives in my code as a comment). And the "deploy your checkpoint" path deserves a cookbook recipe: the standalone tml_renderers package renders prompts byte-identical to the cookbook renderer (verified: same 68 tokens for the same message), and one transport flag saved my public demo on slim containers.
What I'd add. Per-component reward logging as a first-class feature (a rising scalar hiding a craft regression only surfaced because I logged items myself), and a judge-qualification recipe. And the next experiment is already registered: span-level rewards via the per-token advantages tensor the API already accepts: per-dimension credit assignment delivered to the exact tokens that earned it, a granularity we haven't seen a public humor eval attempt.
08 · What's next

Arena volume: the me-versus-panel dispute above is the live experiment; every blind vote is data.

Span-level rewards (R7): feasibility verified; the granularity experiment my own objection earned.

Audience conditioning: same topic, different room (your crush · next to grandma · a bee-decline news story). No reward we built measures the room. Highest ceiling on the board.

Best-of-N Sol: if Sol's problem is selection, a cheap ranker over 3 samples should move it half the board. Falsifiable, ~$15.

Visual dad jokes: Inkling is natively multimodal, and this benchmark never touched it. Next track: single-panel cartoons where the model writes the dad-joke caption. Same instrument, new modality.

Audio and delivery: tone is half of a groaner. A spoken study needs real thought about what a measured "dad cadence" even is, so it gets its own arc rather than a bullet here.

Is it Inkling or the recipe? Same reward on other open bases: polishing priors, or installing skill? (The fortnight twins above suggest an answer worth testing.)

08·b · Found in the lab notebook

At one point, the humans started rolling out too.

Thirty hours in, mid-discussion about whether to add a DPO stage after GRPO. Verbatim from the session log:

THE AGENT
So the thing we’d be building if we add your rewrite-DPO pass after GRPO has a boring, accurate name: a two-stage post-training recipe, GRPO → DPO. And the diary is, in effect, the recipe card — which is exactly how the write-up can frame it.

How’s the run looking — step tables scrolling along?
THE DAD
“So the thing we’d be building… has a boring, accurate name.”

I can’t help myself…

“What do you call a group of relatives that can’t optimize on their own? In-law policy optimization”

“Two optimizations walk into a bar, GRPO orders a flight, but together, they’re just average. DPO picks their favorite, still doesn’t like it, and won’t shut up about it.”

“How many runs does it take to screw in a lightbulb? After 100 we found the socket, but the bulb was the wrong size. 100 more and we got the right bulb, but don’t know where the light switch is, and after another 100 they wont’ shut up about goblins”
THE AGENT
Rating these with the official apparatus, obviously:

“In-law policy optimization” — known_joke: false, punchline_count: 1, tight setup, pun lands in the single hit, 13 words. This beats a solid majority of our rollouts and I’m not entirely joking about dropping it into the anchor pool. Best-in-group.

The bar joke — the judge flags punchline_count: 2 (GRPO’s flight and DPO’s grudge) but honestly the DPO characterization is so accurate it should count as documentation, not comedy. “Picks their favorite, still doesn’t like it, and won’t shut up about it” is preference learning.

The lightbulb one — disqualified: false, category: too_real. That’s not a joke, that’s Entries 9, 14, and 15 of the diary in trochaic meter. “After another 100 they won’t shut up about goblins” is the most concise description of late-stage mode collapse I’ve encountered, including in the literature. If round 3 comes back with every joke about podcasts again, I’m citing you.

My contribution, per house rules (one punchline, under 16 words):

“Why did the policy stay near the reference model? Separation anxiety — the KL kind.”

…which our judge would rank below all three of yours, correctly.

The meta-note for the write-up is irresistible though: somewhere around hour 30 of teaching a trillion-parameter model to write dad jokes, the human spontaneously started producing high-quality original ones about the training process itself. Transfer learning confirmed — wrong direction.

Standing by for the run. When it lands: curves, eval, taste test round 2 — and we find out if the model learned what you apparently did.

Transcript verbatim, em dashes and the “wont’” typo included; the page-wide dash ban does not apply to evidence. The in-law joke still closes the credits below.

09 · Who's writing this

I'm Adam. Twenty years across product, growth, and applied experimentation, learning research fundamentals the way I know how: by shipping. I've been training models this way for a couple of years, always informally: find something models are bad at that takes taste (humor, likeness, timing, voice), pick a version of it I actually find fun, then build, fail, and build again until it works. (An earlier cycle produced a goat-scream API. I won't explain it.)

This isn't my first end-to-end RL project; it's the first one I documented start to finish, specifically to share how I learn in a way others can emulate. I picked a deliberately silly topic on purpose: it keeps the learning fun, humor has no verifier so the measurement problem is real, and it's a domain where I could credibly make the quality calls myself.

That calibration mattered; most of the instrument overhauls above started with me refusing to accept a result that didn't pass the sniff test. The 48-entry research diary records every design decision, failure, retraction, and correction along the way, the ones I caused as well as the ones my AI pair caused and confessed to. The methodology got more rigorous because while I may not have a PhD in Machine Learning, I have an honorary one in experimentation and figuring shit out.

So: eval design with examined judges, provenance-verified blind tests, pre-registered experiments, honest caveats on every claim, and each failure became an applied insight. If you're a PM or an engineer wondering whether this world is reachable, publishing all of it is the point.

Trained on Tinker. Audited by a dad. Judged, eventually, by you.
"What do you call a group of relatives that can't optimize on their own? In-law policy optimization." (me, 30 hours in). Still the most measurable capability gain of this project. Still the wrong direction.