Signal
Loading the stream…
WDSF 2026 results are on record — 10 awards · 11 winnersSee the record →
Loading the stream…
What’s moving in agentic workstations and workflows — drawn from a reviewed source list, scored, and kept at a permanent address you can cite.
Curated and full layers · newest first · scored, sourced, citable
The ledger as a map. Dashed edges are machine-suggested (embedding similarity and duplicate clusters); solid edges are editorial — they appear only where a blog post cites an entry.
Raw data: signal-graph.json
141 entries in the full stream matching the current filters
A peer-reviewed Human Factors study examining how novice programmers' trust in AI-driven Development Environments relates to coding performance and AI compliance when working under time pressure.
Empirical human-factors data on trust calibration in AI-assisted coding; relevant for teams adopting AI dev tools and designing onboarding for junior developers.
A paper proposing a vocabulary for describing multi-agent automated research systems, covering agent identity, operations, communication, visibility, action selection, initialization, and evaluation. It defines a trajectory as a record of one run and distinguishes generative taste (rate of novel proposals) from evaluative taste (gap between proxy score and true quality).
Provides a shared framework for comparing autoresearch system designs and splits the vague complaint about lacking taste into two distinct, testable failure modes.
A study with 24 non-expert users develops a taxonomy of LLM confabulations in immersive 3D scene editing, reports their prevalence and disruptiveness, and defines the perception-reality gap between actual and perceived confabulation occurrences, concluding with design implications for mitigation.
Empirical taxonomy grounded in user observation gives system designers a concrete vocabulary for confabulation types in immersive LLM tools, beyond generic hallucination discussion.
Within-subjects study (n=24) evaluating spatially-anchored graphical previews combined with clarification questions for disambiguating user intent in LLM-assisted parameter-driven geometry editing in VR. The hybrid approach reduced conversation rounds, improved interaction stability, and enhanced user experience versus no disambiguation.
Quantified disambiguation design choices for LLM editing tools in immersive environments, with measured reductions in dialogue overhead and gains in interaction stability. Niche but useful as a reference for VR/AR agent-tool builders.
A study of five expert UI/UX designers using generative AI to design landing-page hero sections finds that design intent is not transmitted via prompts but co-constructed through 'lexical oscillation' between vague, design-domain, and operational language. Mismatches with AI outputs prompt designers to reconsider their intent, shifting engagement from instruction to consultation.
Empirical HCI study documents that prompts are not one-way intent transmission but iterative meaning-making. Useful lens for designers and for builders of generative design tools.
A study with 12 blind and low-vision participants compared two learning formats for complex chart types: (1) tactile chart plus text plus LLM chatbot, and (2) text plus LLM chatbot. Tactile templates supported mental model formation and scaffolded subsequent LLM-mediated data exploration; text-only formats showed weaknesses for spatial-reasoning tasks.
Findings indicate tactile scaffolding improves LLM-assisted learning for non-visual contexts — useful if building AI tools for accessibility or structured data exploration.
Canary is an AI-mediated system for real-time collaborative programming classrooms. It decomposes a peer's programming obstacle into smaller, skill-matched steps and surfaces concrete entry points to potential helpers, aiming to reduce the upfront effort of understanding a teammate's problem and increase peer scaffolding frequency.
Concrete system design for AI-mediated peer help, with formative studies and an evaluation. Relevant to anyone studying how agents structure human collaboration.
Peer-reviewed paper proposing a behavior- and height-based dynamic seat adjustment model that generates coordinated low-frequency seat trajectories aimed at reducing subjective discomfort during prolonged public-transport sitting. The model is preliminarily validated.
Primary-source ergonomics research on dynamic seat motion. Useful as a methodological reference for desk-seat or workstation seat designers, though the study context is public transport rather than office work.
A longitudinal co-design study with five blind and low-vision participants using ProgramAT, an agentic programming tool for camera-based assistive technology. Participants created over 37 custom tools, including some addressing needs unmet by commercial AT. The paper surfaces creation strategies and challenges like model limits and specification conflicts.
Firsthand study of how a specific user group actually uses agentic programming to build personal tools, with concrete recommendations for tool designers supporting non-expert creators.

EvoCode-Bench evaluates coding agents across 227 sequential rounds within a persistent workspace. The analysis finds single-turn scores overstate reliability, with regressions rather than missing features being the primary bottleneck for agent performance.
The sequential-round design surfaces regression behavior that single-turn benchmarks miss, offering a more honest measure of agent reliability for multi-step workflows.
Peer-reviewed paper examining how leaders' role expectations influence their willingness to delegate communication tasks to AI. Authors Raveendhran, Jago, Gratch, and Fast study the psychological factors behind AI adoption for messaging and interpersonal tasks.
Adds a leadership-role lens to AI delegation decisions; useful framing for anyone designing agentic workflows that pass through managerial approval.
A retrieval-augmented, multi-agent LLM framework with human-in-the-loop was tested for detecting cutaneous immune-related adverse events from clinical notes. Compared with unassisted manual review, it raised F1 to 0.88 (vs 0.77), Cohen's kappa to 0.82 (vs 0.50), and roughly halved average review time.
A concrete multi-agent + human-in-the-loop case with measured gains, but confined to a niche clinical NLP task with limited transfer to general agentic workflows.
Formative think-aloud usability study of a chatbot laptop-search advisor with transparency features including constrained generation, on-demand ranking explanations, and comparison. Seven participants completed three tasks. Ease and satisfaction were high, but the ranking explanation caused the most severe usability problem, and several wanted direct-manipulation controls.
Reports a counterintuitive finding that explicit transparency features produced the worst usability problem, with concrete design implications for conversational AI recommenders.
A research paper proposing CRAFT, a smart glasses system that translates real-world experiences into fiction narratives. Through interviews, co-design workshops, and field trials with writers, the study identifies design goals and interaction mechanisms for in-situ creative AI assistance.
Empirical design study with writers on in-situ AI creative support, grounded in three studies rather than speculative vision. Useful as a reference for wearable creative tooling design.
Scoping review of 90 peer-reviewed LLM-based programming support systems in CS education. Introduces the PEA framework (Policy, Enforcement, Authority) for analyzing how systems govern AI assistance, finding varied enforcement mechanisms but centralized authority, and offers a codebook and map of underexplored design configurations.
A structured lens for comparing and designing pedagogically bounded LLM tutoring tools, grounded in a synthesis of 90 systems and explicit about where the design space remains underused.
HARP is a research platform that places participants in controlled scenarios with live, configurable LLM agents. It logs prompt composition behavior (keystroke timing, deletions, pauses), triggers surveys at predefined moments, and allows researchers to control prompts, model parameters, and experimental conditions. The authors illustrate it with a study on how technical specificity and response length affect retention.
A purpose-built testbed for measuring not just what users ask AI but how they compose and revise prompts — a behavioral layer most usability studies miss.
A January 2027 Applied Ergonomics paper (Vol. 138) by Ziang Chen, Zhengyu Tan, and Peiwen Luo uses mixed methods to study how to reduce psychological discomfort that users experience when automation systems produce errors.
Peer-reviewed mixed-methods study on a specific human-automation interaction problem, trust and emotional response to machine failure, that practitioners building agentic systems encounter often but rarely see addressed in the ergonomics literature.

Hands-on review of the Secretlab Magnus Evo XL, a height-adjustable desk with built-in cable management aimed at gaming and home-office use. Assesses stability and value against its price.
Useful buying reference for anyone considering a premium sit-stand desk, with tested notes on stability and cable management rather than spec-sheet marketing.
TargetFinder is a computer vision system using fine-tuned YOLO models for real-time detection of GUI widgets across desktop platforms. Trained on 520 annotated screenshots (Windows, macOS, Ubuntu, web), it outperforms OmniParser and REMAUI baselines and enables system-wide deployment of target-aware pointing techniques like Bubble Cursor and Semantic Pointing. Dataset, models, and library are open-sourced.
Releases a cross-platform widget detection pipeline with dataset and models — a reusable building block for both accessibility pointing aids and agents that need to perceive arbitrary desktop GUIs.

GitHub repository presenting a 'loop engineering workflow' that distills patterns from Anthropic and Google research papers into a structured harness for building with AI agents.
Concrete artifact translating recent frontier-lab agent research into a reusable workflow scaffold. Low HN traction, but the repo itself is the primary source.
Researchers interviewed five blind/low-vision and five sighted scientists about using ChatGPT and Gemini to query figures, diagrams, and tables in scientific papers. They characterized review practices, accessibility workarounds, and conditions causing workflow abandonment, and released a dataset of 115 queries and responses.
Small qualitative study, but the released 115-query dataset and concrete failure modes are usable inputs for designers of AI-assisted technical document readers.
A controlled study with 20 students using a general-purpose AI agent (OpenClaw) across five tasks introduces 'delegation regret' — users regret not the agent's errors but its unauthorized action scope. Trust was calibrated per task; irreversibility combined with external visibility drove trust withdrawal more than stakes alone, and action previews were consistently demanded.
The 'delegation regret' framing and the finding that reversibility-plus-visibility, not stakes alone, drives trust withdrawal are specific design-relevant insights for anyone building or deploying agentic tools.
Scoping review in Ergonomics in Design examining workplace ergonomics in Bangladesh, focusing on anthropometric mismatches between workers and equipment, and associated musculoskeletal disorders amid limited safety implementation.
Surveys ergonomics evidence from a developing-country context often under-represented in mainstream literature, useful baseline for designers and researchers working in similar labor settings.
A peer-reviewed study in Ergonomics in Design evaluating how three seat cushion contours (flat, centre trough, and a third design) affect interface pressure distribution, aiming to identify design strategies that improve seating comfort.
Quantitative pressure-mapping across cushion contours gives readers a measured basis for selecting seating that distributes load more evenly during extended sitting.
Empirical study testing tablet hotspot temperatures (39°C, 43°C, 45°C, 47°C) and ambient temperatures (25°C, 35°C) with 60 participants to examine effects on thermal comfort, device orientation, and use behaviors under a proposed top-center consolidated hotspot architecture.
Quantifies how tablet surface heat changes grip, orientation, and posture — a concrete reference for anyone whose device warms up under sustained use or who specifies tablet hardware.
A peer-reviewed article in Ergonomics in Design arguing that satellite ground station operator interfaces lack sufficient human-centered design and standardization, calling for more research in this area of space operations ergonomics.
A niche industrial-ergonomics call for standardization; relevant to the journal but outside the desk-work focus of the pipeline.
Peer-reviewed study examining how spatial contiguity in visual field and depth affects cognitive load in augmented reality environments used for high-demand multi-screen tasks, published in Ergonomics in Design.
Controlled study on AR layout and cognitive load, useful background for designers planning multi-display or AR workstation configurations.
Exploratory study published in Ergonomics in Design investigating physical discomfort reported by users of Braille displays. The paper examines causes of discomfort during prolonged use of these tactile input/output devices by individuals with low vision or blindness.
Peer-reviewed ergonomics study on a rarely covered input device; useful context for anyone designing or procuring accessible workstations, though relevance to sighted desk workers is indirect.
Three empirical studies applying participatory methods—autophotography, cultural probes, and game-based workshops—to engage employees in workplace research, framed around the rise of hybrid work and the need to capture diverse work-environment experiences.
A peer-reviewed comparison of three participatory approaches for practitioners running employee-centered ergonomics or workplace-design studies, with concrete methodological detail rather than abstract advice.
A peer-reviewed paper in Ergonomics in Design examining how eye-tracking metrics can inform human-centric design of human-machine interfaces in process control operations, framed within Industry 5.0 regulations on operator cognitive and physiological capabilities.
Worth noting for HMI designers and ergonomics researchers in industrial control; documents eye-tracking methodology tied to regulatory context rather than general commentary.
A study in Ergonomics in Design compared trained experts and non-experts using Chile's official ergonomic risk assessment tool for upper-limb injuries. Both groups had difficulty identifying work tasks and consistently rating risks during real-world exercises.
Tests whether Chile's standard ergonomic tool is reliable regardless of practitioner skill. Useful for designers and deployers of risk instruments, since both trained and untrained users struggled similarly.
Academic article arguing that trust in AI is typically examined too narrowly, and proposing more naturalistic methods to study systemic trust as AI takes on high-consequence work. Published in Ergonomics in Design, Vol. 34, Issue 3.
Reframes AI trust as a systemic, naturalistic study target rather than a narrow calibration problem. Useful framing for designers and researchers of agentic systems.
A short feature in Ergonomics in Design arguing AI trust should be modeled on lessons from the modeling and simulation community, shifting emphasis from trust in the model to trust in the result, and highlighting relational over purely technical factors.
Recasts AI trust as a relational, outcome-focused concept rather than a purely technical one, drawing a parallel from established simulation practice.
A Human Factors journal study examining how labeling AI as a teammate, support, or tool affects acceptance through mind perception subdimensions of agency and conscious experience, questioning the premature adoption of the 'teammate' label.
Peer-reviewed evidence on how AI role labels shape user acceptance, directly relevant to designing agentic workflows and team structures.
Peer-reviewed study evaluating how different types of AI explanations affect performance, workload, trust, situation awareness, and user preference in a human-autonomy teaming task within a spaceflight-relevant simulator.
Empirical comparison of XAI explanation styles in a high-stakes teaming context; useful evidence for anyone designing explanations that operators must act on quickly.
Peer-reviewed study in Human Factors examining how the valence and arousal of interruptions affect task resumption and post-interruption performance in younger and middle-aged/older adults under varying task complexity. Finding: affective interruptions impaired task performance less than neutral ones across age and task demands.
Counters the blanket assumption that all interruptions hurt equally. The valence of the interrupting content matters, and the effect holds across age groups and task loads — relevant for anyone structuring knowledge work.
A Human Factors paper proposing a structural model of how attentional effort varies over time during a vigilance task, linking effort dynamics to subjectively inferred context, along with an estimation method and empirical test.
A formal model of vigilance effort over time is useful background for anyone designing or scheduling sustained-attention monitoring tasks at a desk.
Human Factors journal study examining how direct versus indirect automated cues in a decision support system influence visual search performance and attentional allocation, with implications for safety- and health-critical domains.
Empirical comparison of cue types in decision support, useful for designers of monitoring interfaces and anyone studying how automation steers attention.
A study in Human Factors examining how the moral intensity of a decision and the perceived trustworthiness of an AI certifier influence individuals' willingness to revise their decisions when given AI-generated suggestions in workplace contexts.
Empirical evidence on which conditions make workers accept or override AI recommendations — useful input for designing agent-assisted workflows and calibrating human-AI trust.
A Human Factors journal study examining how task priority and task difficulty, and their interaction, influence task-switching decisions when operators manage multiple supervisory tasks distributed across displays.
Primary peer-reviewed evidence on task-switching across screens, useful for anyone structuring multi-display workflows or designing supervisory interfaces.
A Human Factors journal paper comparing system-wide trust and component-specific trust perspectives in multi-agent human-agent teams, examining individual variability in trust bias across different agent teammates.
Relevant for designers and researchers of multi-agent systems who must decide whether users evaluate agents as one unit or separately. Adds empirical nuance to a debated theoretical framing.
Peer-reviewed replication and extension examining how response-effect, stimulus-response, and stimulus-effect compatibility jointly influence performance in lever tool use, testing whether effects observed with simple button responses generalize to lever inputs.
Tests whether compatibility effects from button-pressing studies hold for lever inputs, a useful baseline for designing or evaluating non-keyboard input devices.
Human Factors journal study modeling expertise in manual sanding, identifying sanding strategies that correlate with improved performance outcomes and reduced physical stress during the task.
Empirical motor-skill data on which sanding strategies reduce strain; of interest to those training manual finishing work or studying physical ergonomics of skilled tasks.
A peer-reviewed paper in Human Factors proposing Relative Advantage Theory to explain decision-making mechanisms in emergencies where humans collaborate with AI, accompanied by initial empirical evidence supporting its assumptions.
Provides a grounded theoretical lens for when humans should defer to or override AI in time-critical scenarios, relevant to designers of safety-critical human-AI systems.
Peer-reviewed study examining how AI-provided explanations affect efficiency, diagnostic accuracy, user perceptions, and workflow integration in ophthalmologists' clinical diagnostic and treatment workflows, identifying challenges in human-AI collaboration.
Empirical findings on the efficiency costs of explainable AI in real clinical decision-making. Useful beyond medicine for anyone designing human-AI workflows where explanations add cognitive overhead.
An article in Human Factors examines methodological considerations for conducting vigilance research in nontraditional, non-laboratory settings, outlining challenges and offering practical guidance for researchers designing field studies.
Useful for HFE researchers planning field vigilance work; clarifies trade-offs between experimental control and ecological validity that affect how lab findings transfer to real desk-work contexts.
Peer-reviewed Human Factors study showing that post-task self-reports of mind-wandering shift when forced errors are introduced, suggesting retrospective thought-content reports may reflect perceived performance rather than causally explain it.
A methodological finding with direct bearing on how we interpret attention and distraction self-reports in knowledge work and workflow research.
A study in Human Factors examining how user exposure to an automated decision support system and the saliency of its errors affect estimates of automation reliability, evaluating sensitivity and calibration of those judgments.
Quantifies how perception of automation reliability shifts with experience and error visibility — useful background for designing or evaluating trust in AI-assisted workflows.
Controlled study comparing paper-based, user-fixed AR, and world-fixed AR assembly manuals on task performance, dorsolateral prefrontal cortex hemodynamic responses, and perceived workload. Published in Human Factors, Vol 68, Issue 9.
fNIRS-measured cognitive load across manual formats gives workplace designers workload data rather than preference surveys, useful for AR rollout decisions.
Peer-reviewed Human Factors study comparing a large set of specific job rotation schemes, measuring biomechanical risk, body discomfort, and psychosocial demands to assess effectiveness at reducing musculoskeletal disorder risk among workers.
Empirically compares many rotation schemes rather than endorsing rotation as a blanket control, offering evidence for designing rotation cycles in desk-based and mixed-task work.