Real products
teach AI
what people value.
AI Matey is my applied research studio. I turn real products and real user behavior into evals, rewards, and training signal that product and business teams can actually use.
Why
I came to AI research through product, growth, marketing, and creative experimentation. Twenty years of finding cracks in systems, forming hypotheses, testing them against reality, and using behavior to make the system better.
Model training and reinforcement learning turned out to be the same craft with the difficulty turned up. The thing being improved is no longer a funnel or a campaign. It is behavior.
Thesis
The next wave of AI will not be shaped by bigger models alone. It will be shaped by better definitions of success, better feedback, and better judgment about what people actually value. The richest source of all three is real products and real experiences.
Products are not just apps. They are where people reveal value: edits, rejections, retries, purchases, saves, corrections, shares, and moments of trust. I turn those signals into evals, reward surfaces, and training loops, and bring that value to the product and business teams building with AI.
The goal is not better AI in the abstract. It is AI that learns what useful means in the real world, measured on the real thing, not a toy.
Field report
DadBench: I RL-trained a trillion-parameter model to tell dad jokes, then built a benchmark rigorous enough to measure whether it actually worked. A working case study in the hard part: designing rewards and evaluations for something subjective without getting Goodharted, the same problem every real product team runs into.
Experiments
- SmilePlease: AI portraits for likeness, trust, and identity preservation.
- Foodways: culturally grounded recipe and food-history RAG and deep research loops.
- Appree: communication rewriting without factual drift.
- GoatScreams API: subjective audio generation for timing, humor, and vibe.
In progress
- Slay the Sprint: a roguelike PM deckbuilder for RL training.
- IntroGym: a multi-agent, human-and-agent training loop.
- LinkedinCRO: career and aspiration signals for LinkedIn optimization.
Field reports, as they ship
DadBench is the first in a series. I document each project end to end, so you can see what applied research actually looks like from the inside: the rewards that got gamed, the judges that got fired, the findings that held up.
New write-ups go out through Context Drift, my newsletter on how people actually adopt AI.
Subscribe on Context Drift →