Research

Four threads. They are separable on a whiteboard and not separable in practice — each one is mostly a way of making the other three cheaper.

Thread — Proposal

Hypotheses worth an experiment

A model that emits a thousand ideas has not helped anyone. The useful problem is ranking: which conjectures earn a test, and whether a system can be calibrated about the novelty of its own output rather than uniformly confident about all of it.

Thread — Verification

The half of the loop that can say no

A claim counts when something outside the model can reject it. We build that half — simulators, proof checkers, retrieval against prior art, and handoffs to instruments and wet labs that return a real measurement. Generation is the part everyone shows you. Rejection is the part that makes it science.

Thread — Search

The cost of being wrong, per finding

Search has a budget, measured in reagents, compute, and months. Most of our engineering is not about proposing better hypotheses but about lowering the number of failed experiments it takes to reach one that holds.

Thread — Transfer

Whether any of this transfers

It would be convenient if discovery were one skill. We do not assume it. We run the same system against unrelated search spaces — catalysts, theorems, protocols — and report where the transfer breaks, including when the answer is unflattering to the method.

What we cannot answer yet

Published here because a lab that only lists what it has solved is telling you half of its work. If you have spent time on any of these, we would rather hear from you than from someone who agrees with us.

  1. How do you measure novelty without a human in the loop, and without rewarding a model for being merely unusual?
  2. Can a system recognise that its own hypothesis is unfalsifiable before it spends a month on it?
  3. What is the right unit of credit when a model and a person co-author a finding neither would have reached alone?
  4. Does a discovery method transfer across fields, or is every field its own search space with its own geometry?
  5. What should a lab refuse to search for, and who decides that while the search is still expensive?

Notes

Short pieces on method and results, negative results included. The full posts go up shortly.

The loop these threads run on

OpenRecursive is where all four threads meet a real search space. Code, run logs, and rejected candidates are public.