Real products
teach AI
what people value.
AI Matey is my applied research studio. I do lab-grade training work (environments, reward design, and evals), and I know how to make it land with real customers and real workflows at scale, not just benchmarks.
Why
I came to AI research through product, growth, marketing, and creative experimentation. Twenty years of finding cracks in systems, forming hypotheses, testing them against reality, and using behavior to make the system better.
Model training and reinforcement learning turned out to be the same craft with the difficulty turned up. The thing being improved is no longer a funnel or a campaign. It is behavior.
I did not learn RL and then go looking for a problem. I spent two decades inside real products, learning what people actually value, and why. Then I picked up the training stack to act on it. That is what makes training land.
Thesis
The next wave of AI will not be shaped by bigger models alone. It will be shaped by better definitions of success, better feedback, and better judgment about what people actually value. The richest source of all three is real products and real experiences.
Products are not just apps. They are where people reveal value: edits, rejections, retries, purchases, saves, corrections, shares, and moments of trust. I turn those signals into evals, reward surfaces, and training loops.
The goal is not better AI in the abstract. It is AI that learns what useful means in the real world, measured on the real thing, not a toy.
I am that bridge. Frontier labs know how to train models. The harder, less glamorous part is connecting that training to real customers, real workflows, and the tacit knowledge people carry that never survives into a benchmark. That connection is the work I do, and it is what AI Matey is for.
Field report
DadBench: I RL-trained a trillion-parameter model to tell dad jokes, then built a benchmark rigorous enough to measure whether it actually worked. A working case study in the hard part: designing rewards and evaluations for something subjective without getting Goodharted, the same problem every real product team runs into.
What I build
I build training systems around problems where the hard part is defining and teaching good behavior, not just generating the next token.
Training environments and agent gyms
Game-like and simulated environments for teaching models long-horizon decision-making, adaptation, memory, resource allocation, subjective judgment, and multi-agent behavior. Recent work includes a roguelike product-management environment, adversarial creative competitions, and environments built around delayed rewards and hidden objectives.
Real work → training signal
Systems that turn human expertise, workflows, interviews, and changing real-world jobs into structured tasks, evaluations, reward signals, and training environments. This includes tooling for generating environments at scale and continuously updating them as the underlying work changes.
Product-grounded evals and rewards
Experiments in translating fuzzy human preferences into measurable model behavior: humor and taste, likeness and identity preservation, factual fidelity in rewriting, culturally grounded research, and subjective audio quality. The common problem is figuring out what good means, measuring it without getting gamed, and turning that judgment into something models can learn from.
Selected systems
- DadBench · Subjective preference RL, reward design, model evals · Field report →
- Slay the Sprint · Long-horizon agent training through simulated product decision-making.
- Environment infrastructure · Modular tooling for multi-agent environments, shared state, rewards, and domain-specific simulations.
- Expertise → environments · Pipelines that transform real workflows and expert knowledge into scalable training tasks and evals.
Earlier product experiments have explored training models for likeness and identity preservation, culturally grounded research, factual fidelity in rewriting, multimodal taste, and other forms of human preference.
Get in touch
Have a messy behavior you need a model to learn? Turning it into an environment, reward, and evaluation is the work I do. Email me.
Field reports, as they ship
DadBench is the first in a series. I document each project end to end, so you can see what applied research actually looks like from the inside: the rewards that got gamed, the judges that got fired, the findings that held up.
New write-ups go out through Context Drift, my newsletter on how people actually adopt AI.
Subscribe on Context Drift →