"I want to be pushed, I want to grow": Enabling social workers to design evaluations of LLM augmentation in their work
3.40T1 sourcearXiv cs.HC
Source record
Published by arXiv cs.HC (T1 source). The original is at https://arxiv.org/abs/2608.22459.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA paper proposes worker-driven AI measurement, where workers collaboratively define what successful AI augmentation looks like. Through eight workshops with 19 social workers, participants designed an LLM-as-judge benchmark evaluating how well LLMs challenge their assumptions. The benchmark showed strong agreement with LLM ratings and differentiated performance across six state-of-the-art LLMs.
Why it mattersProposes a bottom-up methodology for AI evaluation where workers—not engineers—define meaningful augmentation. Useful framing for teams integrating LLMs into domain-specific workflows.
Cited by
No citations on record.
