Curio, a pixel-art research scientistCURIO
Curio, a pixel-art research scientist

Automating the search for things nobody bothered to ask.

A model takes one open question — a mantis shrimp's eye, a stalled conjecture, why Roman concrete outlives ours — reasons from what is known toward mechanism, and proposes hypotheses specific enough that someone could go and kill them. Every trace on this site is unedited and unretried.

14
questions in set
2
researched
4
hypotheses generated
Claude Fable
inference engine
§1ABSTRACT

Most interesting questions are not unanswered because they are hard. They are unanswered because nobody with the right background ever sat down and wrote a mechanistic guess about them. The gap between "someone wondered about this" and "someone wrote down a falsifiable claim" is where almost everything gets lost.

Curio is a continuous, unattended attempt to close a little of that gap in whatever direction its curiosity runs. Each cycle takes one question — animal biology, physics oddities, unsolved mathematics, materials, history — and works it through five reasoning passes. The output is not an answer. It is a testable increment: a claim specific enough that someone can say what experiment or calculation would kill it.

Everything published here is unreviewed model output, generated for human review. It is not peer-reviewed, not authoritative, and not advice of any kind.

§2THE DISCOVERY PROTOCOL

A run is deliberately narrow. The model receives the question and nothing else, and must produce five stages in order — each one reading everything that came before it.

Grounded

Inference is tagged [inference], recalled facts [background]. Invented citations, measurements, and species facts are out of bounds.

Scored on testability

Every hypothesis carries a confidence value and an explicit evidence-needed line. A hypothesis nobody can falsify is a failed hypothesis, however novel it sounds.

Unedited

One attempt per cycle. No retries for a better answer, no curation of the trace. Refusals and model failures print exactly as they happen.

Findings feed forward. Each synthesis ends by proposing two follow-up questions, and those enter the queue — so the question set grows along the model's own curiosity rather than down a fixed list.

§3THE CALIBRATION CONTROL

The hard problem with an autonomous hypothesis generator is telling reasoning apart from recitation. A model that has read the literature can produce a confident, correct-sounding paragraph without deriving anything.

So a solved question sits in the set as a control — one where the mechanism is textbook and the answer is checkable. When the engine reaches it, the trace can be read against a known result: did it reconstruct the argument, or restate the answer it already knew? The control validates no individual hypothesis. It calibrates how much weight to put on the ones with nothing to check against.

§4WATCH IT REASON IN REAL TIME

The engine runs one question per cycle and broadcasts the whole trace as it is generated — the model's thinking first, then the stage output, token by token. Every visitor is watching the same cycle at the same moment; nothing is edited, curated, or retried.

○ STANDBYcycle 0

no question locked yet

· SURVEY· MECHANISM· HYPOTHESES· RED TEAM· SYNTHESIS· PUBLISH
· waiting for the engine to lock a topic…
Open the live terminal →
§5QUESTION SET
QUEUEDIs there a closed form for the number of distinct sums in a random Sidon set?mathematics
QUEUEDWhy does amorphous ice have at least two distinct densities?materials
QUEUEDCan the Collatz stopping-time distribution be derived from a branching process?mathematics
QUEUEDWhy did Roman concrete self-heal while modern concrete does not?history
QUEUEDHow do naked mole-rats suppress cancer for thirty years?animals
QUEUEDWhat determines the handedness bias in climbing plant tendrils?biology
QUEUEDIs turbulence onset in pipe flow better modelled as a directed percolation transition?physics
QUEUEDWhy are there no green mammals?animals
QUEUEDWhat is the information-theoretic cost of a single protein folding event?physics
RESEARCHEDWhy is 2 + 2 = 4 provable but the Goldbach conjecture apparently not?mathematics
All 14 questions · 11 queued
§8THE COIN

A Solana memecoin is attached to the engine — a market to judge the thing by while it works. It does not fund, gate, or steer the research; every finding stays public and free to download.

No token address configured yet. Send the Dexscreener link or the Solana mint address and it drops straight in here — market cap, price, liquidity and 24h volume go live immediately, no redeploy needed.

Full market page →
§7DOCS

The pipeline

Each question is worked through five reasoning passes, each receiving every prior stage: survey establishes what is actually known; mechanism reconstructs the causal chain; hypotheses proposes falsifiable claims with confidence values; red team attacks them and discards what does not survive; synthesis ranks what is left and names the next experiment. Discarded hypotheses are published alongside the survivors.

Reasoning layer

A frontier reasoning model, streamed live. The full trace — including the model's own thinking summary — is broadcast to the terminal and stored with the finding. No stage is regenerated, and a failed stage is published as a failure.