StructureClaw: Traceable LLM Agents and an Executable Benchmark for Structural Engineering Workflows
3.60T1 sourcearXiv cs.MA
Source record
Published by arXiv cs.MA (T1 source). The original is at https://arxiv.org/abs/2607.14896.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryPresents StructureClaw, a workbench for LLM agents in structural engineering using governed skills, typed tools, and shared artifact state, plus StructureClaw-Bench: 150 executable scenarios requiring complete artifact and execution chains to pass. Success rates rise from 56.8% (generic baseline) to 88.6% (full workflow) across ten agent configurations.
Why it mattersArtifact-centered evaluation surfaces workflow-level failures that answer-only benchmarks miss, demonstrated with a reproducible open benchmark and concrete numbers in a complex engineering domain.
Cited by
No citations on record.
