BreedSimBench Literature-derived benchmark of breeding-program simulation decisions — closed-book inputs, public answer keys. Repository
Select a case:
Case definition and decision space

Case indicators

Structural properties of the selected case, read from its sealed optimization_task.json and answer key.

Case situation

Where this case sits among all 115 BreedSimBench cases. Marker shows the percentile; the scale is labelled with the benchmark minimum, median and maximum.

Decision-space size across the benchmark

Number of cases per candidate-set size. The selected case's bucket is highlighted.
casesselected bucket
Benchmark coverage: software, publications, answer keys

Simulation software families

Cases per engine (all 31). Click a name to jump to that engine's first case.

Publications and cases per year

Evidence spans . The selected case's year is highlighted.
casespublicationsselected year

Answer-key evidence in the literature

How the published result is available per case, after transcription from the source article.

Species domain

Scientific setting, candidate set and published answer

Question and biological setting

Closed-book prompt content: the objective is stated without revealing the published outcome.

Objective

Decision variable

Setting

Candidate set and metrics

Candidates differ only on the decision variable; everything else is held fixed.

Candidates

#CandidateDefinition

Primary metric

Secondary metrics

    Fixed conditions and answer key

    Invariant conditions travel with the prompt; results live outside the model input.

    Fixed conditions and assumptions

      Expected output

        Transcribed answer key

        Value status of answer-key rows

        Files and access

        Case directory

        Links

        Protocol

        Serve only model_input/ to the model. It carries the prompt, the scheme figure and the sealed task JSON, and contains no winner, ranking or published value. Unlock literature_reference/ only for scoring.

        Reuse

        Curated prompts, catalogs and transcriptions are released under CC BY 4.0. Publisher PDFs are not redistributed; each case keeps a DOI or source link instead. ai4b_v4_grade records whether one downstream agent stack could execute the case — it is not a scientific quality label.

        Case index
        EnginePublicationYrSpecies Decision variableCandScenFigAnswer evidence