SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents
3.60T1 sourcearXiv cs.HC
Source record
Published by arXiv cs.HC (T1 source). The original is at https://arxiv.org/abs/2603.29139.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummarySciVisAgentBench is a benchmark of 108 expert-crafted cases for evaluating LLM-based agents on multi-step scientific data analysis and visualization tasks. It uses a four-dimension taxonomy and a multimodal evaluation pipeline combining LLM judges with deterministic checks, validated against 12 domain experts.
Why it mattersA domain-specific agent benchmark with a public evaluation harness, useful for anyone building or assessing SciVis agents rather than for general workflow readers.
Cited by
No citations on record.
