Real products
teach AI
what people value.

AI Matey is my applied research studio. I do lab-grade training work (environments, reward design, and evals), and I know how to make it land with real customers and real workflows at scale, not just benchmarks.

Why

I came to AI research through product, growth, marketing, and creative experimentation. Twenty years of finding cracks in systems, forming hypotheses, testing them against reality, and using behavior to make the system better.

Model training and reinforcement learning turned out to be the same craft with the difficulty turned up. The thing being improved is no longer a funnel or a campaign. It is behavior.

I did not learn RL and then go looking for a problem. I spent two decades inside real products, learning what people actually value, and why. Then I picked up the training stack to act on it. That is what makes training land.

Thesis

The next wave of AI will not be shaped by bigger models alone. It will be shaped by better definitions of success, better feedback, and better judgment about what people actually value. The richest source of all three is real products and real experiences.

Products are not just apps. They are where people reveal value: edits, rejections, retries, purchases, saves, corrections, shares, and moments of trust. I turn those signals into evals, reward surfaces, and training loops.

The goal is not better AI in the abstract. It is AI that learns what useful means in the real world, measured on the real thing, not a toy.

I am that bridge. Frontier labs know how to train models. The harder, less glamorous part is connecting that training to real customers, real workflows, and the tacit knowledge people carry that never survives into a benchmark. That connection is the work I do, and it is what AI Matey is for.

Field report

DadBench: I RL-trained a trillion-parameter model to tell dad jokes, then built a benchmark rigorous enough to measure whether it actually worked. A working case study in the hard part: designing rewards and evaluations for something subjective without getting Goodharted, the same problem every real product team runs into.

Read the field report and play the arena →

What I build

I build training systems around problems where the hard part is defining and teaching good behavior, not just generating the next token.

Training environments and agent gyms

Game-like and simulated environments for teaching models long-horizon decision-making, adaptation, memory, resource allocation, subjective judgment, and multi-agent behavior. Recent work includes a roguelike product-management environment, adversarial creative competitions, and environments built around delayed rewards and hidden objectives.

Real work → training signal

Systems that turn human expertise, workflows, interviews, and changing real-world jobs into structured tasks, evaluations, reward signals, and training environments. This includes tooling for generating environments at scale and continuously updating them as the underlying work changes.

Product-grounded evals and rewards

Experiments in translating fuzzy human preferences into measurable model behavior: humor and taste, likeness and identity preservation, factual fidelity in rewriting, culturally grounded research, and subjective audio quality. The common problem is figuring out what good means, measuring it without getting gamed, and turning that judgment into something models can learn from.

Selected systems

Earlier product experiments have explored training models for likeness and identity preservation, culturally grounded research, factual fidelity in rewriting, multimodal taste, and other forms of human preference.

Get in touch

Have a messy behavior you need a model to learn? Turning it into an environment, reward, and evaluation is the work I do. Email me.

Field reports, as they ship

DadBench is the first in a series. I document each project end to end, so you can see what applied research actually looks like from the inside: the rewards that got gamed, the judges that got fired, the findings that held up.

New write-ups go out through Context Drift, my newsletter on how people actually adopt AI.

Subscribe on Context Drift →