
A model takes one open question — a mantis shrimp's eye, a stalled conjecture, why Roman concrete outlives ours — reasons from what is known toward mechanism, and proposes hypotheses specific enough that someone could go and kill them. Every trace on this site is unedited and unretried.
Most interesting questions are not unanswered because they are hard. They are unanswered because nobody with the right background ever sat down and wrote a mechanistic guess about them. The gap between "someone wondered about this" and "someone wrote down a falsifiable claim" is where almost everything gets lost.
Curio is a continuous, unattended attempt to close a little of that gap in whatever direction its curiosity runs. Each cycle takes one question — animal biology, physics oddities, unsolved mathematics, materials, history — and works it through five reasoning passes. The output is not an answer. It is a testable increment: a claim specific enough that someone can say what experiment or calculation would kill it.
Everything published here is unreviewed model output, generated for human review. It is not peer-reviewed, not authoritative, and not advice of any kind.
A run is deliberately narrow. The model receives the question and nothing else, and must produce five stages in order — each one reading everything that came before it.
Inference is tagged [inference], recalled facts [background]. Invented citations, measurements, and species facts are out of bounds.
Every hypothesis carries a confidence value and an explicit evidence-needed line. A hypothesis nobody can falsify is a failed hypothesis, however novel it sounds.
One attempt per cycle. No retries for a better answer, no curation of the trace. Refusals and model failures print exactly as they happen.
Findings feed forward. Each synthesis ends by proposing two follow-up questions, and those enter the queue — so the question set grows along the model's own curiosity rather than down a fixed list.
The hard problem with an autonomous hypothesis generator is telling reasoning apart from recitation. A model that has read the literature can produce a confident, correct-sounding paragraph without deriving anything.
So a solved question sits in the set as a control — one where the mechanism is textbook and the answer is checkable. When the engine reaches it, the trace can be read against a known result: did it reconstruct the argument, or restate the answer it already knew? The control validates no individual hypothesis. It calibrates how much weight to put on the ones with nothing to check against.
The engine runs one question per cycle and broadcasts the whole trace as it is generated — the model's thinking first, then the stage output, token by token. Every visitor is watching the same cycle at the same moment; nothing is edited, curated, or retried.
no question locked yet
The engine is not pointed at one field. Click an area to see the questions inside it and what is being worked on right now.
3 questions · 1 researched · 1 queued
Is there a closed form for the number of distinct sums in a random Sidon set?
RESEARCHING NOW
1 questions · 0 researched · 1 queued
Why does amorphous ice have at least two distinct densities?
OPEN AREA
1 questions · 0 researched · 1 queued
Why did Roman concrete self-heal while modern concrete does not?
OPEN AREA
6 questions · 1 researched · 5 queued
How do naked mole-rats suppress cancer for thirty years?
OPEN AREA
1 questions · 0 researched · 1 queued
What determines the handedness bias in climbing plant tendrils?
OPEN AREA
2 questions · 0 researched · 2 queued
Is turbulence onset in pipe flow better modelled as a directed percolation transition?
OPEN AREA
A Solana memecoin is attached to the engine — a market to judge the thing by while it works. It does not fund, gate, or steer the research; every finding stays public and free to download.
No token address configured yet. Send the Dexscreener link or the Solana mint address and it drops straight in here — market cap, price, liquidity and 24h volume go live immediately, no redeploy needed.
Each question is worked through five reasoning passes, each receiving every prior stage: survey establishes what is actually known; mechanism reconstructs the causal chain; hypotheses proposes falsifiable claims with confidence values; red team attacks them and discards what does not survive; synthesis ranks what is left and names the next experiment. Discarded hypotheses are published alongside the survivors.
A frontier reasoning model, streamed live. The full trace — including the model's own thinking summary — is broadcast to the terminal and stored with the finding. No stage is regenerated, and a failed stage is published as a failure.