FR

About

Theo Martin, portrait

Nine years putting ML and then LLMs into production, first on pricing at Amazon, then on product catalogs in SaaS. Today I use coding agents every day, at a rate of hundreds of billions of tokens, and I work with teams that want more out of them: developers, product managers, designers.

The path

The same journey, in pixels, an animation that moves scene by scene as you scroll.

I started at Amazon, in Luxembourg: three and a half years on European pricing. I came in through supply chain optimization and then AWS architecture, with four AWS certifications passed in four months, including the Solutions Architect Professional. in pixels

Then pricing, as the only data scientist on a team that went from three people to thirteen. Causal inference, with synthetic control coded by hand before the libraries matured, hierarchical Bayesian price elasticity models, and hedonic prices estimated on text and image embeddings, with ELMo and then fine-tuned BERT in 2019. in pixels

Out of that came the price volatility monitoring program, which I started and which ended up adopted across the group worldwide. Thirty billion price changes analyzed, an anomaly caught across a billion visits. in pixels

Then Unifai, where I was the company’s first ML engineer. An end to end MLOps pipeline on GCP for retail groups, and language models from before ChatGPT, FLAN-T5 zero-shot against my fine-tuned extractor, on real industrial catalogs, better on booleans, worse on numbers. That work is what put the company in a position to be acquired: during due diligence, I walked investors through the architecture and ML infrastructure, through to the Akeneo acquisition in 2023. in pixels

Then two and a half years at Akeneo as Tech Lead of the Core AI team. I architected and ran the internal inference platform every product team depended on, and behind them hundreds of enterprise retailers. I wrote the library that sends that traffic to VertexAI, OpenAI and Anthropic through a single LiteLLM proxy, with each tenant’s cost tracked in the response headers, and per-request log volume dropped by half to three quarters. And the Data Architect Agent, a multi-agent system with human gates that took catalog onboarding for an enterprise retailer from several months down to a few days. in pixels

From the tool to the harness

At Akeneo I pushed Cursor into my team, then Claude Code as soon as it shipped. I gave classes to several teams, and people I had convinced then convinced their own teams in turn. I also pushed the leadership, early and hard, to pay for licences: in the end everyone got a Claude Code seat. in pixels

Then I moved from the tool to the harness: CLAUDE.md files versioned in the repo, so everyone’s harness gets better without each person having to look after their own. My setup ended up as the basis for an internal training.

What I think

Positions reviewed on October 5, 2026. They move with the tools, so the date counts as much as the sentence.

The first win stops many teams

I have never yet seen a team where the tool was the limiting factor. You install the agent, it writes a few tests, they pass, and that first win convinces everyone the tool is under control. That day it had read the right files; the next day it goes another way, a script, a grep, the shell, part of the rules stop applying, and nobody sees it. That is usually where things freeze. The same mechanism works the other way round: one failed first try is enough to convince people they have understood that it does not work.

What I am aiming at is a team that delegates for real, not one that has the agent write a few tests once and stops there. The gain comes back as ambition, as going after things you would not have let yourself start before. Anthropic describes its own internal usage in those terms, down to letting Claude write a whole feature in an autonomous loop from Figma mockups. That matches what I see, without proving it, and I keep in mind that Anthropic sells the tool, publishes its wins and burns tokens without counting. I have no financial tie to Anthropic, Codex or Cursor, and my advice holds for all three.

Harnesses are roughly equivalent today

That is not “the tool doesn’t matter”, which is false, and which sounds like a salesman talking down what you already know so he can talk up what he is pushing. It is a dated observation about the market: the harnesses have converged, so the gap has moved elsewhere, to what you put inside them. The CLAUDE.md files, the hooks, the skills, the structure of the repo, what the agent can check on its own before handing back to you.

If a tool takes a clear lead again in six months, the sentence falls and I will say so.

An achievement is worth mostly its date

Describing what I built is not worth much on its own any more. Automatic PR review by an agent is something GitHub itself commoditized in April 2025, a checkbox in a branch ruleset. Presenting it today as a feat means presenting yourself as someone who has just discovered the subject.

So the part of my work that mattered most is not the review tooling, which GitHub has made ordinary since. It is CLAUDE.md files versioned across a whole team. Those files exist everywhere, AGENTS.md is in more than 60,000 projects according to agents.md, and what is rare is that they hold up over time, as the note below shows.

In late August 2026 I read 269 public repos file by file, not through a code search that rounds its own counts, to see who actually owns the files that steer coding agents. One organization out of the 269 had automated the expiry of its own, with a named owner per file and a cron, on its internal mirror, that opens a ticket when the date passes. Seven of its entries were already overdue the day I looked. I am saying that from the outside and from public sources, the list is not published and the figure dates from late August, so if your team does better, write to me, that interests me more than the opposite.

What disappears is a trade, not people

In September 2024 I wrote in a short post that “80%+ of us” would be obsolete, “probably within 3 years”. I was wrong on one point: it is not developers who disappear, it is the part of their work that consists of turning a clear specification into code. The trade moves towards specification, verification and decision, and the specification itself moves up a level when you let the agent look for its own context (Notion, Confluence, Jira, Slack) instead of believing you can hand it everything. I do not believe in a handmade niche: whoever buys software does not look at how it is made, unlike a vase or a piece of blown glass. The sentence falls if postings for developers without AI rebound. The figures I have seen show a market closing at the entry level rather than a disappearance.

We became reviewers, and reviewing is a probability

My job looks like that of a director checking on his seniors: result-oriented, with tests the agent has to pass to prove itself. I agree with what Karpathy says about orchestrating agents. Reviewing lowers the probability of error without eradicating it, five reviews beat two, never zero errors. With robust tests on isolated parts, some zones stop being reviewed, except what is hard to test (race conditions) or critical. The sentence falls if a measurement shows that code never reviewed but well tested breaks more in production.

I test everything, my own harness included

On my repo, a CLAUDE.md cut from 108 to 30 lines gave 22 cases out of 22 before and after on my bench, and 2 out of 16 with no file at all. The bench is small and its judge is a simple rule. I give it as proof that you can measure this at home in an hour, not as a general result. The sentence falls if an independent measurement shows that a longer file holds up better.

What I am working on

I have worked with large companies, one of them in the CAC 40. At a large company, what blocks agent adoption is almost never technical: it is who decides, on what criteria, and when someone writes the decision down somewhere.

I take short engagements, often remote, for teams that want their harness to pay off. An audit to find out whether your guards hold when the agent switches tools, a workshop day inside your repo with the work committed that same evening, regular office hours to catch what drifts, or a build when what you need is shipped code. The audit is the simplest way in.

Services and prices

I also write now and then about what I break and what I fix along the way.

Read the writing


Based in Paris, EU timezone, overlap with US East until 5pm ET. Remote anywhere, on site anywhere in France. contacttheomartin@gmail.com, LinkedIn, GitHub.