Get started
Start with the Baseline pilot: export tasks, run your agent, review the evidence, and summarize the results.
Set up the benchmark.
Install the evaluation package in a Python environment. Python 3.10 or later is required.
git clone https://github.com/BOBQWERA/BioinfoAgentBench.git
cd BioinfoAgentBench
python -m venv .venv
source .venv/bin/activate
python -m pip install -e .Export the pilot tasks.
Export the instructions for the defined 100-task Baseline pilot.
bioinfoagentbench tasks --tier baseline --pilot \
--output task_inputs.jsonl
mkdir -p runTask definitions are available in the repository. Baseline and Fingerprint input data, and Baseline reference outputs, are supplied separately. All Frontier case inputs are included.
Run your agent and review its outputs.
Give the agent each instruction and its input files. Save the outputs under run/, then review them against the task’s verification rules. Record one JSON object per task and agent in judgments.jsonl.
Each passing check requires an existing evidence file and a matching quote or a location description. Cover every verification rule and all 100 pilot tasks.
View an example judgment
Example after verifying that the output for BL-0251 reports the required value:
{
"task_id": "BL-0251",
"agent": "my-agent",
"all_rules_reviewed": true,
"checks": [{
"criterion": "Median treeness difference: fungi minus animals",
"passed": true,
"reason": "The reported difference is 0.05.",
"evidence": [{
"path": "BL-0251/answer.txt",
"quote": "0.05"
}]
}]
}Score and summarize.
The CLI validates the supplied judgments and evidence references, then aggregates success and available resource metrics.
bioinfoagentbench score --tier baseline --pilot \
--records judgments.jsonl --evidence-root run/ \
--output scored.jsonl
bioinfoagentbench summarize --tier baseline --pilot \
--records scored.jsonl --output summary.json