Innovation is search, and measurements decide the game

What would it take to build a machine that does science on its own? Not a tool a researcher picks up, but a system that runs the loop itself, proposing experiments, reading what comes back, and deciding what to try next, then getting better at all of it as it goes. I’ve come to believe that scientific research and engineering design are the same kind of activity wearing different clothes. Both are search over an action space. The moves in that search are what I’ll call unit operations, small trustworthy steps you can compose, a term I’m borrowing from chemical engineering. Once you see research this way, a lot of things that felt like separate skills start to look like moves in one game: you sketch a design, identify the observables, run experiments, rank the results, and iterate. None of this is new, exactly.1 What makes it worth revisiting now is auto-research. If scientific research really is a search over unit operations, then that is also how you would build the machine: a harness that runs the search itself and grows its own operations.

What you actually search over

A research question never arrives alone. It comes with an environment: the system you’re poking at, the instruments you have, the things you’re allowed to vary, the way a fluid mechanics problem comes with a control volume you draw around it. You start by flailing, crude experiments to see what the environment does when you push on it, and out of the mess you pin down two things: the action space, the moves you can make, and the observables, what you can actually measure.2

Those moves are the unit operations, the smallest reusable steps that each reliably do one thing: a measurement you know how to take, a transformation you know how to apply, a sub-problem you know how to solve. Decompose a problem into them and science becomes searchable. You stop asking how to make something work and start asking which sequence of known moves gets you there. In a tool like CAD 3 the unit operations were curated for you; in science oftentimes they aren’t handed over, and finding and trusting the set of unit operations is most of the work of scientific research.

Where innovation actually comes from

Innovation runs along a spectrum, and most of its mass sits at one end.

At the common end is recombination. You take operations that already exist and find a new sequence, a new arrangement, that hits an objective no one had aimed at before. This is most of engineering, and it’s genuinely valuable; the moves were lying around, but nobody had composed them that way. One of my favorite examples is adaptive optics. Astronomers built it to un-blur stars: a deformable mirror, a sensor that measures how the atmosphere has scrambled the incoming light, and a fast computer that bends the mirror to undo the damage. Microscopists then borrowed the very same setup to un-blur cells buried deep in tissue, where the sample scrambles the light the way the atmosphere does. Same operations, a new target, a sharper view of the very small. Nobody had to invent a new primitive. They pointed the old ones somewhere new.

At the rare end is inventing a new unit operation, a move that wasn’t in anyone’s repertoire: a new measurement, a new transformation, a new primitive. This expands the search space instead of exploring it, and it’s what might be called true scientific innovation. It’s rare, and the line is blurry: most “new” operations turn out on a second look to be clever recombinations. The genuine ones are the handful you read about in history: the laws of motion, and the calculus that grew up alongside them. They feel like they came from nowhere. Almost everything else, when you look closely, came from somewhere.

Then there’s what you can measure

Recombination and new primitives are both about the moves: finding them, or inventing them. There is a third lever, closely related but worth separating: changing the coordinates of the search by changing what can be measured. A new observable is a new axis on the system, one that makes legible some structure that was there in the material all along. For example, stiffness and toughness belong to the same composite, but until fracture mechanics handed you toughness as a number, there was no axis to aim at. The unit operations haven’t really changed; the space gained a dimension because you learned to measure along it.

A search can only climb a signal it can sense, and that same fact bites the moment you build a search to run on its own. An agent that measures only stiffness will never find toughness, however good its exploration, because the axis it would travel doesn’t exist for it. So a harness makes a quiet third choice, next to the moves and the objective: what the agent is allowed to measure. For a coding agent the observables are its tests and metrics; for an experimentalist in the wet lab, the observables are often confined by the instruments. With wrong observables, the research would end up optimizing hard along the few axes you happened to wire up.

Identifying good test cases is itself costly. At the easy end a unit test runs in milliseconds and the harness gets a clean signal on every move, which is the regime where coding agents thrive. It climbs fast. A large codebase that keeps evolving is already hard to test, and you often can’t tell whether a change helped without a full training run that burns days and a pile of compute. Science sits further along still: a simulation can run for days, an experiment for months, and for a genuinely new observable there might be no immediate test to write at all.

Handing it to a machine

The payoff of seeing scientific research as a search over unit operations is that it makes a question operational: the operations you trust, the observables you choose, the tests that certify them, and the search over their compositions all become things you can actually build. The verifier is a component like the others, and usually the most expensive one to get right.

Concretely, the unit operations stop being a metaphor and become things an agent can call: a CLI command, a REST endpoint, an MCP server, a sandbox service with a clean signature. The version everyone actually wants goes further: a harness that grows its own, proposing new operations and new observables and folding the good ones back into the set it searches over. That is self-evolution, and the same catch holds. It can propose as fast as you like, but it moves only as fast as it can certify what it proposed.

Choosing the environment, identifying the observable, picking the objective: these sit upstream of the search. A better search won’t conjure what isn’t there, and it won’t fix a badly drawn environment or a wrong observable. But if recombination really is where most innovations live, then a great many of them are already sitting in that space, in arrangements no one has tried yet. That’s exactly the region an automated search can reach.

It also says where the machine wins first. The obvious half is that search pays off earliest where operations are thick and answers are cheap to check. The half worth saying out loud is that “cheap to check” is usually an observable in disguise: a field automates once it has the right thing to measure. And when a field is stuck, it’s more often for want of the right observables than for want of operations. Most research isn’t waiting on a new idea from nowhere, it’s waiting to become searchable.

  1. Herbert Simon called design a search through problem spaces more than half a century ago, and Brian Arthur describes innovation as the recombination of what already exists. ↩

  2. Here’s the loop in miniature, from a problem a materials engineer might actually face. You’re designing a particle-reinforced composite, a soft matrix stiffened by hard particles. You get to choose the particle size, the volume fraction, and the arrangement, though not every choice is reachable: some you can’t fabricate cleanly, and some are too fine or too dense for a direct simulation to handle. So you map that practical boundary, then sweep what’s left, hunting for a target you can measure like the effective stiffness. That’s search in its plainest form. The operations are fixed; you’re recombining them to maximize a number. ↩

  3. Think about CAD. The shape of a part is, in principle, an infinite-dimensional object; there are uncountably many surfaces you could draw, and no software could ever search that. So every CAD program quietly refuses to. Instead it hands you a finite set of operations, extrude, revolve, fillet, sweep, boolean, and lets you build the shape as a sequence of them. The infinite space of geometry collapses into a finite space of operation sequences, and that you can navigate, edit, and optimize. Those operations are the unit operations, and the parametric model is your search space. ↩