Work is changing. Your space is how you answer.
Following that change — through real workspaces, live signals, and a yearly festival.
WDSF 2026 — winners on record
Off-season · See the 2026 resultsDesk Setup of the YearChen Sifan · Hangzhou, China · S·0068
Why this ledger exists
- 01
Judgment and taste
When AI does more of the doing, the human part of work gets sharper — judgment, taste, direction.
- 02
An observation post
AI is rewriting the workday in real time. Beyond Desk watches where that change becomes physical.
- 03
Authorship
How you work is becoming something you design, not something you are given.
Setups. 178 real workspaces, logged as submitted — people first, gear second.
browse all →
S·0057Student遇到困难睡大觉 · Sydney, Australia
S·0121Software developerMars Xiang · Shanghai, China
S·0051Flight attendant一只狗 · Shanghai, China
S·0052Abdullah bin Mohammad · Saudi Arabia
S·0053PhotographerAkayu · Guangzhou, China
S·0054EducatorAlana (The Simple Norm) · United States
S·0055Software EngineerAlex · Boca Raton, FL
S·0056Product designerAlex Richard · Dongguan, China
Scan. The one tool here that reads your own desk — a private AI report, by email.
scan your desk →One photo, read as a working system — scored 1–5 on the same four dimensions the festival uses, each with a written reason.
- ReportArrives by email — no total, no ranking
- PrivacyPrivate to you — kept out of the gallery and the registry
- StorageHeld in access-controlled storage, then deleted on a fixed schedule
The four dimensions it reads
- 01Work-mode fitDoes the layout serve a believable, describable workflow?
- 02Spatial narrativeCan a stranger read the person and place from the frame?
- 03Craft & executionHow completely is the intent finished — not how much it cost?
- 04AuthenticityA space in real use, or staged for show?
Blog. Thinking built on the evidence — essays and notes reasoned from recorded signals.
all entries →- B·0005
Safe models do not compose into safe systems
Six papers in two weeks converge on one engineering fact: safety measured on a single agent does not survive wiring agents together. The failure lives in the joints — memory, context, setup files, and the management layer itself.
- B·0006
Delegation regret is measurable now
Three studies put structure on a feeling agent users know well. The precise version: people regret the scope an agent took, not the errors it made.
- B·0004
This week in the record: agents are leaving the IDE
Six recorded entries from one fortnight point the same direction: the coding agent's home is no longer the editor buffer. It is Slack, Linear, the issue tracker, the phone.
Evidence graph
The ledger as a map. Dashed edges are machine-suggested (embedding similarity and duplicate clusters); solid edges are editorial — they appear only where a blog post cites an entry.
- cites — editorial citation (blog post → entry)
- related — machine-suggested (embedding similarity)
- cluster — same story, archived duplicate
AI hardware & peripherals
- Adjust your Navigator Trackball angle and sensitivity
- Press to exit auto-mouse and other Navigator firmware updates
- iFLYTEK P1 Pro: The AI-powered recorder slash lighter
- DIY Voyager Angled Keycaps
- Layout Buffet - Backlighting
- INNOCN CB32U1 review: An affordable alternative to the Apple Studio Display?
- Flagship convertible with Snapdragon X2 Elite - Microsoft Surface Pro OLED 2026 Review
- Excellent business laptop with 64 GB RAM - Lenovo ThinkPad T14s Gen 7 AMD Review
- June Roundup of Navigator Updates
- Layout Buffet - Hyper, Meh, and Software Shortcuts
- DIY Navigator Moonlander Shells
- DIY Moonlander Wrist Rest Pad
- DIY Navigator Thumb Module
- DIY Navigator Trackpad Shells
- iFLYTEK AI Recorder P1: An AI assistant as a wearable device
- AI at the Edge is a different operating environment
- HP Reveals Keyboard Computer with Ryzen AI Chip
- Keychron's Nape Pro turns your keyboard into a laptop‑style trackball rig
- Exploring Different Keyboard Sensing Technologies
- Toucan Wireless Split Keyboard with Touchpad
- OpenAI and Broadcom unveil LLM-optimized inference chip
- 192GB of VRAM in One PC… The Cheap Way
- AMD Built the DGX Spark Rival I Predicted… But There's a Catch
- Building the World’s Smallest AI Workstation
- Most Powerful 16" Gaming Laptop Has a Secret
- NVIDIA RTX Spark Made Everyone Mad
- 3 New PCs, One Giant AI Model… This Shouldn’t Work
- Three months wrong about why my 4-node AMD cluster was slow
- The Non-NVIDIA AI Card Everyone’s Ignoring
- INSANE $1000 32GB VRAM Local Ai Server
- $1500 Local AI Server Build Tested with Hermes Agent Gemma 4 and Qwen 3.6
- This $129 AI Storage Server Upgrade Saved Me $2,800
- BEST Local Ai Motherboard CPU Combo You've Never Heard Of
- KEEP AI LOCAL! Explaining Agentic AI and The Loop: FT MSI Cubi NUC+ and the MSI EdgeXpert Mini PCs
- Is STRIX Better than SPARK? Now Launching w/new Software: AMD's Ryzen AI Halo Developer Workstation
- NVIDIA's Secret Windows Upgrades for N1/N1X Laptops Help Everyone; AMD Could Benefit the Most
- Minisforum's N5 Max is an Absurd NAS
- Analyzing Nvidia GB10's GPU
- Inside Nvidia GB10’s Memory Subsystem, from the CPU Side
- Nvidia’s B200: Keeping the CUDA Juggernaut Rolling ft. Verda (formerly DataCrunch)
- Today was the perfect day for Poolside to drop Laguna S 2.1 because I just got these in! Finally have a half decent amount of VRAM. 3x V620 = 96 GB.
- RTX 5090 and 5080 Ran the Same Local AI Until the VRAM Ran Out
- Got these baddies in the mail today (2X 3080 20GB)
- Apple’s Hidden AI Model… The Speed they never showed
- Apple M5 isn't making full use of its matmul cores yet
- How much are RTX PRO 6000s going for in your country/state?
- PSA: DO NOT use Intel consumer platforms for multi-GPU setups
- Building the World’s Smallest AI Workstation
- Taking a Look at Gigabyte's MW94-RP0 Xeon 6 E-ATX Motherboard
- Is it worth getting 128GB MacBook Pro? Will it ever be comparable to today’s frontier models for coding?
- AMD Says 2 Ryzen AI Halos Can Run a 400B Model... I Tested It
- World's First(?) Underwhelming AMD Ryzen AI Halo Cluster
- AMD Strix Halo with 32 GB RAM for $1899 - Asus ProArt PX13 2026 Convertible Review
- A user has managed to run Kimi K3 on 80xRTX 5090, via 25GbE Ethernet.
- My second Inspur AGX-2 with another x8 v100 arrived!
- PCIe Gen6 and Gen5 Will Both Matter for AI Storage
- DGX Station running GLM5.2 in VSCode
- ASUS Showcases NUC 16 Family Powered By Panther Lake
- "Data center in a Box (on Wheels)" 256Gb VRAM/512Gb RAM AI Server 6-8 Month Operational Review, Stability Write Up, Benchmarks
- Checking Out The GMKTek X3 Strix Halo: More Strix Halo Shenanigans!
- SK hynix, In Collaboration With SanDisk, Unveils The New High Bandwidth Flash (HBF) Standard, Helping To Resolve AI Inference Bottlenecks, Targeting Up To 3TB/s Bandwidth
- Kimi K3 full model running on 16x GB10 cluster at 20+tps
- Powerful creator tablet with Snapdragon X2 Elite - Asus ProArt PZ14 Review
- Get AI max+ 395 laptop or wait for rtx spark?
- Google DeepMind reshuffle 🧠, Meta Muse Code 💻, Anthropic chip team 🧩
- We Upgraded Our ZOTAC 4090 24G to Have More VRAM!
- RTX 5090 Owner Built An Open-Source Tool That Shuts Down PC If It Detects The 12VHPWR Cable Drawing Too Much Power, But It Can Only Work On Specific GPUs
- A 10M IOPS Kioxia GP1 SSD Shown Running at FMS 2026
- Showoff Saturday: Local 4x 6000 Pro (multi-year progression)
- enabling PCI-E p2p for consumer Nvidia cards will yield you more than you think
- RTX 5090 96GB spotted on Alibaba?
- Underestimated budget solution: radeon 780m iGPU
- WisdPi WP-UT9 USB 10GbE Adapter Review
- Comu Action Pro: AI Assistant put to the test
- Comu Action Pro: AI assistant put to the test
- Show HN: A tiny LLM running at 21,000 tok/s on a $250 FPGA (Live Demo)
- Minisforum N5 Max Review with AMD Ryzen AI Max+ 395
- I built a weird, low-power llama.cpp server using an Intel N100 + RTX 5060Ti
- Ling-3.0-flash quant ladder on one DGX Spark: the whole thing sits in a 32 to 40 tok/s band
- Low Power AI Is More Efficient than NVIDIA // The AI Hardware Show S2E10
- OdinLake O3 WireControl ergonomic chair review: Better than the LiberNovo Omni?
- Nvidia doubles RTX PRO 6000 Blackwell's MSRP to a staggering $16,000 — 96GB card started pre-orders below $8,000 last year
- You could purchase a Desktop with 2TB of DDR5 - It only sets you back some $200k+
- 27-inch monitor with a 240 Hz IPS panel, G-Sync, and accurate colors: KTC H27E6S review
- M5 MacBook Air Against Every Generation for Dev Work
- This was a data center a year ago… Now it's on my desk
- This 13-in-1 USB-C docking station with built-in GaN+ charger changed my charging habits for the better — a review
- club-5060ti refresh: tested RTX 5060 Ti presets, a proper high-context harness, and Qwen3.8 27B
- MOKiN's 13-in-1 USB-C docking station has an actually useful gimmick and it turned my Lenovo Legion Go into a usable desktop PC
- Show-off Saturday: Intel Arc B140 build.
- The dream is to reach 200GB VRAM
- Qwen3.8-27B Q8_0 on Strix Halo is seriously impressive
- CDW has bumped the MSRP of the RTX Pro 6000 from $16,000 to $19,999
- Linux Improves VRAM Management in 7.3 Kernel 🥳
- This laptop with 64 GB RAM is perfect for AI agents: HP ZBook 8 G2a 14 review
- Alibaba's RISC-V CPU, XuanTie C950, Runs Qwen-3.8 27B at 30 tps
- Framework 13 Pro...Why Apple Solders Memory
- A Server with EIGHT GPUs? ft. ASRock R9600D
- Buying a V100/older NVIDIA GPU? Run this to check for older memory issues
- “The All Spark” Cluster: Upgrading from 16 - 36 DGX Sparks
- GMKtec is going to launch new hardware with Ryzen AI Max+ PRO 495 at IFA Berlin 2026
- With AI lumbar support, massage, heating, and cooling: Hbada X7 office chair review
- Apple M5 Server
- Apple introduces new Mac Studio with M5 Max and M5 Ultra - up to 512GB of unified memory
- Apple releases M5 ultra at 1.2TB/s bandwith
- Intel Arc Pro B60 Dual 48G spotted
- Mac Studio M5 Max Cost Analysis
- Mac Studio M5 Ultra is a DGX Spark Killer for Local AI
- 224GB of GPU Memory on 1 Desk and It Should Not Work
- EliteBook X Flip G2i 14 AI review: HP's strongest convertible yet is also one of the priciest
- [AINews] Hot Chips: OpenAI’s Jalapeño, Cerebras CS-5, Groq 3 LPX, Apple M6
- Let's Talk about SR-IOV and Proxmox on Intel Arc B60 and B70!
- I Built a Monster PC Just to Run AI Locally
- ROCm 10.0: A Decade of Open Compute, Built for the Age of Agentic AI
- Hot Chips 2026: Samsung’s Processing-in-Memory (PIM)
- Exo labs claiming 4.8 tb/s memory bandwidth through m5u Mac Studio clustering
- DGX Spark cluster
- It's official! 192GB Framework
- Ran Qwen3.8-Flash-Next (79 GB, 2-bit) at 350K ctx for 3.5 hours on a 128 GB M5 Max — speed vs context depth, 100 turns, one graph
- When you say, because I can. Limits of X870e
- Experience report - Qwen 3.8 Flash Next on memory rich, GPU poor setup
- NVIDIA® DGX Station™ Delivering Data-Center-Class Performance from the Desktop
- This $60,000 Mac Cluster Has a $10 Problem
- HP ZBook Ultra G1a 14 with Strix Halo review: A victim of the memory crisis?
- A very confusing report from Puget Systems
- HiDock P1 AI voice recorder hands-on review
- 2/5 of my CMP 170HX have died after 2 weeks and the 3rd came with defective tensor cores. Current prices DO NOT justify the risk you are taking
Agentic workflow patterns
- A couple of days ago, I sat down with Vivek Bharathi and dumped my brains. Here's the interview...
- a sneak preview behind an embedded software factory. I suspect rapid application dev is back
- porting software has been trivial for a while now. here’s how you do it.
- don’t waste your back pressure
- teleporting into the future and robbing yourself of retirement projects
- everything is a ralph loop
- i ran Claude in a loop for three months, and it created a genz programming language called cursed
- Note #728
- The twilight of the chatbots
- Claude Dispatch and the Power of Interfaces
- Three Years from GPT-3 to Gemini 3
- The Shape of AI: Jaggedness, Bottlenecks and Salients
- Management as AI superpower
- A Guide to Which AI to Use in the Agentic Era
- On Working with Wizards
- Real AI Agents and Real Work
- An Opinionated Guide to Using AI Right Now
- Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO
- [AINews] Lilian Weng summarizes 35 papers on Harness Engineering for RSI
- Vercel's Andrew Qu on why agents are a new kind of software
- The website of the future may assemble itself for every visitor
- Skill engineering and the case against one-shot AI design
- How Cursor deploys AI inside the enterprise
- Warp CEO Zach Lloyd on why software factories are the next phase of coding
- Autoresearch: The feedback loop behind self-improving agents
- Forward Deployed Engineers and the future of software engineering
- AIEWF Daily Dispatch: Loops, Software Factories & Forward Deployed Engineers
- Co-Existence and the End of Co-Intelligence
- CogniConsole: Externalizing Inference-Time Control as a Formal Abstraction for Reliable LLM Interactions
- Scaling Managed Agents: Decoupling the brain from the hands
- How we contain Claude across products
- Building a C compiler with a team of parallel Claudes
- How we built Claude Code auto mode: a safer way to skip permissions
- Harness design for long-running application development
- Demystifying evals for AI agents
- Effective harnesses for long-running agents
- Code execution with MCP: Building more efficient agents
- How we built our multi-agent research system
- Effective context engineering for AI agents
- Claude Code: Best practices for agentic coding
- Building effective agents
- The "think" tool: Enabling Claude to stop and think in complex tool use situations
- Tuning the harness, not the model: a Nemotron 3 Ultra playbook
- Your coding agent bill doubled. Here’s how to fix it.
- Why Model Neutrality Matters More Than Cloud Neutrality
- Improving Agents is a Data Mining Problem
- Wiki Memory
- How to Use RLMs in Deep Agents
- How Candidly Built State-Aware Agent Harnesses with LangSmith
- Introducing Dynamic Subagents in Deep Agents
- Building Durable AI Agents
- AIUC-1: Building trust in AI agents
- Zero Trust for AI Agents
- Humility in the Age of Agentic Coding
- Agentic Coding and the Economics of Open Source
- Post-Mortem of Anthropic's Claude Code Leak
- Inside an AI-Run Company
- Beyond chatbots: Agents that tackle your SOPs
- While loops with tool calls
- Mosaic: Runtime-Efficient Multi-Agent Embodied Planning
- When is Routing Meaningful? Diversity and Robustness in Language Model Societies
- Secret Scanner Agent: Extracting Secrets and Access Context from Unstructured Documents
- Shared Selective Persistent Memory for Agentic LLM Systems
- Leveraging Multi-Agent System (MAS) and Fine-Tuned Small Language Models (SLMs) for Automated Telecom Network Troubleshooting
- Rlm-Workflow
- Show HN: Durable AI agents without the workflow engine
- Claude and Codex and Grok: my current workflow and its friction
- Show HN: Sigil – FIDO2 key-derived P2P remote desktop (agentic workflow retro)
- "Code Is Cheap. Show Me the Talk.": Lessons from Teaching and Managing AI Coding Tool Usage in a Visualization Course
- Intervenability as a Design Requirement for Autonomy and Oversight within Human-Centered AI
- Spatula: Exploring On-Demand In-Situ Interfaces and Interaction for Attribute Control
- Motif: Discovering and Automating Personal Web Workflows
- U-Lens: Supporting User Uncertainty Management in Long-Form LLM Responses
- Memory-Conditioned Tool Calling for Camera-First Visual Agents
- Exploring Agentic Workflows for Generating High Quality Math Visual Aids
- Neutralizing Structural Inequality in the Nigerian FinTech Sector
- Distributed Agent System: Fault-Tolerant Collaboration Among Embodied Agents
- Auditing Belief-Conditioned LLM Agents in Hidden-Information Social Deduction Games
- An Explainable Agentic System for Detection of Conversational Scams with Summary-Based Memory
- Multi-Agent LLMs Fail to Explore Each Other
- How Much Does Correctness Cost? Budgeted Placement of Strong Correctors in a Weak Multi-Agent Swarm
- Closed-Loop Control with Rule-Aligned Small Language Models and Multi-Agent Self-Correction
- Replicating Belief, Not Bits: Epistemic State Replication for Agentic Systems
- Agentic Context Learning with Self-Discovered Specification
- Verification of Adaptive Agentic Controllers through Finite Rule Revision
- Norm Enforcement for AI Agents: Robustly Shaping Behavior in Multi-Agent Systems
- Automated Textbook Auditing with Multi-Agent LLM Systems
- Who&When Pro: Can LLMs Really Attribute Failures in AI Agents?
- Can Agentic Trading Systems Pay for Their Own Intelligence?
- Can LLMs Perform Deep Technical Comprehension of Computer Architecture Papers?
- StructAgent: Harness Long-horizon Digital Agents with Unified Causal Structure
- When Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent Systems
- 5 Trends That Defined AI Engineering at World’s Fair 2026
- Compos3D: Interactive Part-Based Composition for Creative Control in Generative 3D Models
- A\"ira: Rethinking AI Research Assistants for Interdisciplinary Science
- Designing Agent-Ready Websites for AI Web Agents: A Framework for Machine Readability, Actionability, and Decision Reliability
- SheetMind: An End-to-End LLM-Powered Multi-Agent Framework for Spreadsheet Automation
- Self-Regulated Reading with AI Support: An Eight-Week Study with Students
- ParaTutor: Coordinating Parent and Child Math Tutoring through Role Separated LLM Scaffolding
- MetaInfer: A Knowledge Only LLM Inference Engine Generator SKILL Toolbox
- Graph Feedback Controls Consensus and Clique Formation in Open-Weight Language-Model Populations
- RCWT: Measuring Task-Budget Displacement from Coordination Content in LLM Calls
- Internet of Agentic Things: Networked AI Agents for Closed-Loop IoT Orchestration
- XScientist: A Git-Like Research Protocol for Long-Running Autonomous Scientific Discovery
- Agent Identity URI Scheme: Topology-Independent Naming and Capability-Based Discovery for Multi-Agent Systems
- Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents
- Networked Intelligence: Active Shared Context Graphs for Human-AI Team Science
- Learning Engagement Assistant (LEA): Cross-Course Scalability and Classroom Evaluation of an Agentic AI Tutoring System
- Persona Migration and Expectation Recalibration in Generative AI Adoption: A Longitudinal Study at a State Department of Transportation
- When Bots Join the Team: Bot Adoption and the Institutional Fabric of Open-Source Software Projects
- Learning Latency-Aware Orchestration for Multi-Agent Systems
- Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems
- Benefits and Limitations of Communication in Multi-Agent Reasoning
- MASPRM: Multi-Agent System Process Reward Model
- Project Kaleidoscope: Contextual, Human-Aligned Evaluation for Real-World AI Applications
- Setup Complete, Now You Are Compromised: Weaponizing Setup Instructions Against AI Coding Agents — cited by 1
- Beyond Interestingness: Semantic and Context-Aware Natural Language Query Recommendations for Visual Data Analysis
- Proving the ROI of agentic AI in financial services
- Latent Communication Between Language Model Agents: Channels, Alignment, and the Limits of Text
- Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation
- ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System
- Orchestrating Power Grid Studies with Multi-Agent AI and MCP Servers
- AI Agents Do Not Fail Alone:The Context Fails First — cited by 1
- When Is Delegated Play Truthful? Within-Range Regret and the Trilemma of Aligned Delegation — cited by 1
- Towards an Intention Abstraction Layer for Autonomous Industrial Systems
- Bad Memory: Evaluating Prompt Injection Risks from Memory in Agentic Systems — cited by 1
- StructureClaw: Traceable LLM Agents and an Executable Benchmark for Structural Engineering Workflows
- Does Multi-Agent Debate Improve AI Feedback on Research Papers?
- Proposed spec to share SKILL and "loop"/"Workflow" by OCI registry
- Ask HN: Workflow Automation vs AI Agents?
- Let AI take over your API debugging workflow
- Dealing with increasingly complicated agents
- We've all done RAG, now what?
- How Cars24 scales conversations and builds faster with OpenAI
- How to manage AI investments in the agentic era
- How sales teams use ChatGPT Work
- How data science teams use ChatGPT Work
- Australian Payments Plus moves faster with ChatGPT and Codex
- How agents are transforming work
- Codex-maxxing for long-running work
- Automating the Rubber Stamp: What If an Agent Ran Your Deployment Gate?
- Show HN: I RL-trained an agent that trains models with RL (for ~$1.3k)
- The Pulse: What can we learn from Bun’s rapid Rust rewrite with AI?
- Context engineering with Dex Horthy
- What is “loop engineering?”
- Impressions from visiting OpenAI, Anthropic, & Cursor
- The Pulse: a trend of trying to cut back on AI spend within eng departments?
- Ideas: slow down to speed up when working with AI agents
- What building Shippy taught us about building agents
- Model Routing Is Simple. Until It Isn’t.
- Shipping huggingface_hub every week with AI, open tools, and a human in the loop
- MosaicLeaks: Can your research agent keep a secret?
- Is it agentic enough? Benchmarking open models on your own tooling
- How an Agent Built a 3D Paris Gallery by Chaining Two Hugging Face Spaces
- I was giving my coding agent context the wrong way...
- How to build proactive agents & self-improving company (Fully explained)
- Ralph-loop 2.0? The real autonomous coder is coming...
- 3 New PCs, One Giant AI Model… This Shouldn’t Work
- Breakdowns for Human-Machine Creative Reflexivity
- Perceived AGI: Believability as Dimensional Completeness, Not Capability
- When Not to Automate: A Formal Protocol for Human Preservation in AI-Optimized Organizations
- HiLSVA: Design and Evaluation of a Human-in-the-Loop Agentic System for Scientific Visualization
- Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation — cited by 1
- CoWeaver: A Bi-directional, Learnable and Explainable Matching Engine for Mixed Human-Agent Science Collaboration
- The Honest Quorum Problem: Epistemic Byzantine Fault Tolerance for Agentic Infrastructure — cited by 1
- Crayotter: Traceable Multi-Agent Workflows for Long-Form Video Editing
- A hierarchical memory architecture overcomes context limits in long-horizon multi-agent computational modeling
- Tmux + Fable = Cut 35% less token
- Building Governed Agents: A Framework for Cost, Control, and Compliance
- TaskArtisan: Designing Composable Generative Widgets for LLM-Assisted Analysis
- Sidekick: Designing Communication for Effective Multitasking with Computer Use Agents — cited by 1
- Real-World Evaluation of an AI Agent Drafting Translational Impact Summaries
- Can LLM Code Explanations Adapt to Diverse Problem-Solvers' Needs?
- Large Language Models in Architecture Studio: A Framework for Learning Outcomes
- Automated Hardware Validation Test Plan Generation for Large Scale AI Datacenter Platforms Using a Generative AI Multi-Agents Architecture
- Auto Research for Materials: Auditable AI-Scientist Workflows with Held-Out Transfer
- Clarify Before Executing: A Self-Evolving Agent for Resolving Intent Asymmetry in 3D Tool Orchestration
- RELIC: Revealed Principles for Learning Interpretable Composable Skills in Multi-Agent Planning
- Agentic ERP: Multi-Agent Large Language Model Architecture for Autonomous Enterprise Resource Planning
- Autonomous Discovery of Wireless Communications Algorithms
- Towards Agentic Agent-based Models: Feasibility, Performance, and Statistical Model Checking
- MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for Parameter-Efficient Reasoning in Compact Models
- FinSAgent: Corpus-Aligned Multi-Agent RAG Framework for Evidence-Grounded SEC Filing Question Answering
- O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning
- OMAC: A Holistic Optimization Framework for LLM-Based Multi-Agent Collaboration
- FOCAL: Filtered On-device Continuous Activity Logging for Efficient Personal Desktop Summarization
- ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog
- How ZigZag made our workflow worse before we fixed it
- Ask HN: I stopped fighting AI over-reliance and built a workflow around it
- Workflow time travel to prevent agent Vision Drift
- The Archaeologist’s Copilot
- DSLs Enable Reliable Use of LLMs
- Fragments: July 13
- Experiences with local models for coding
- Fragments: July 6
- Building Reliable Agentic AI Systems
- Fragments: June 16
- Fragments: June 2
- Fragments: May 27
- The test suite as a regression sensor
- The VibeSec Reckoning
- Bliki: Vibe Coding
- Fragments: May 14
- Bliki: Interrogatory LLM
- What is Code
- Fragments: May 5
- Fragments: April 29
- Structured-Prompt-Driven Development (SPDD)
- Fragments: April 21
- Feedback Flywheel
- Harness engineering for coding agent users
- Encoding Team Standards
- Do Automated Evals Work?
- “It’s Hard to Eval” Is a Product Smell
- The Revenge of the Data Scientist
- Why I Stopped Using nbdev
- FORGET Loop Engineering. Agentic Engineering is about THIS
- Claude Fable 5 BANNED: The First Model Agentic Engineers DON'T NEED
- Pi Coding Agent Observability: HTML Specs with Gemini 3.5 Flash and GPT Image 2
- Top #1 Opportunity for Senior Engineers: Agentic Engineering
- Pi to Pi: Two-Way Agent Orchestration with the Pi Coding Agent
- GPT-5.5 VERIFIED Opus 4.7: A Pi Coding Agent That REVIEWS Like YOU
- MAXIMIZE Your Claude Code Subscription (Without Getting BANNED)
- How Agents Manage Other Agents: Four Subagents Patterns in 2026
- How Autoresearch will change Small Language Models adoption
- Agents: Inner Loop vs Outer Loop
- Can We Close the Loop in 2026?
- The Agent Client Protocol Overview
- The importance of Agent Harness in 2026
- Context Engineering for AI Agents: Part 2
- Why (Senior) Engineers Struggle to Build AI Agents
- Agents 2.0: From Shallow Loops to Deep Agents
- The Rise of Subagents
- Pydantic AI 2.0: The New Best Way to Build AI Agents is Composing Capabilities
- The Best AI Coding Setup Isn't the Most Autonomous One (Here's Why)
- Finally, an Open Standard for the Karpathy LLM Wiki is HERE
- Google Just Dropped a Masterclass on Agentic Engineering (It's SO Good)
- The Creators of Claude Code and OpenClaw don't Prompt Their Agents Anymore?!
- Patterns for Building Cybersecurity Evals
- Using LLMs to Secure Source Code
- How to Work and Compound with AI
- Product Evals in Three Simple Steps
- Software Factories, Light and Dark
- Own the Outer Loop
- Earning taste and judgment
- Agentic Autonomy Levels
- The New Software Lifecycle
- Agentic Code Review
- Loop Engineering
- The Intent Debt
- The Orchestration Tax
- Understanding is the new bottleneck
- Code like a surgeon
- On owning a codebase, and why it may be the hardest job in software
- Automating Security Triage with HackerOne and Deep Search
- How we're using Sourcegraph and a Slack bot to detect vulnerabilities and react quickly
- Why coding agents fail in large codebases (and what to do about it)
- Building DataBot: Our always-on data assistant
- The Coming Loop
- SEPs Are Moving to Pull Requests
- The Coding Agent Is Dead
- Go Deep
- Tab, Tab, Dead
- Handoff, Please
- Stick a Fork in It, It's Done
- Software Is Made Between Commits
- Introducing Zed's Agent Metrics
- On Programming with Agents
- AI's 70% Problem
- MAIS: Exploring human-AI interaction in fair and transparent recruitment with a multi-agent LLM-based system
- Asymmetric Encounters with AI: Professional Designers’ Perceptions and Integration of AI Tools
- Crafting AI Explanations for Every Role in Your Enterprise
- Vibe Architects: Agentic Vibe Coders
- The Core Skill of Design in the AI Era: Critique
- Context Architecture
- Embracing AI with Claude's C Compiler
- Fragments: July 21
- How Apollo Rebuilt Its AI Assistant on Deep Agents to Power the Full GTM Loop
- Assistant or Actor? Student Trust, Control, and Delegation Regret When Using a General-Purpose AI Agent — cited by 1
- AInimation: Animating from Prompt to AI-Generated Responses
- EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration
- HALO: Interactive Co-abductive Reasoning in Scientific Hypothesis Generation
- A Decision-Centered Reference Architecture for Trustworthy Agentic Commerce
- Engineering Trustworthy Agentic AI for Critical Systems
- Solve the CyberGym benchmark
- 3 Years of Graph Engineering with LangGraph
- Multiplayer
- Building an AI-orchestrated publishing workflow for a long-form writing project
- Agentic Workflow's Cache Keepalive Costs 8x Too Much
- Loop Engineering Workflow Based on Anthropic and Google Papers
- NTT DATA Group cuts incident analysis to 30 minutes with Codex
- How to Actually Run Your Coding Agent Safely (And Avoid the Horror Stories)
- A Framework of User Experience Principles for Human-AI Agent Interaction in the Workplace
- Not Birds of a Feather: Personality-Based Partner Selection in LLM Agents
- Knowledge-Centric Self-Improvement
- Show HN: Hanesu – An experimental workflow layer for AI coding agents
- Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models
- Copilot cloud agent for Linear is now generally available — cited by 1
- Exploring the Design Space of LLM-Based Programming Support in CS Education: A Scoping Review through the Lens of Assistance Governance
- Thinkink: 2D Spatial Ink-native Interaction with LLMs
- pAI-Econ-claude: A Gated Human-in-the-Loop Multi-Agent Architecture for AI-Assisted Economic Theory Development
- FedAgentKE: Federated Semantic Knowledge Evolution for Heterogeneous Agents
- ChannelGuard: Safe Models Do Not Compose into Safe Multi-Agent Systems — cited by 1
- Human-in-the-Loop Large Language Model Framework for Identification of Cutaneous Immune-Related Adverse Events
- Workload-Aware Caching for Multi-Agent Systems
- MKEvolve: A Modular Multi-Agent Framework for Kernel Code Generation
- UX-Context Design: Using UX Knowledge to Inform AI-Generated Design
- How much are you actually using your local models these days? Which ones do you reach for the most?
- What does it mean to "own your intelligence"?
- Deepseek V4 flash - Hy3 or is Qwen3.6 27B still the most solid for agentic/coding?
- Control panels to clarify user intent with Large Language Models
- Reliability-Contagion Feasibility in LLM Multi-Agent Networks
- When Language Models Meet NeuroGraphs: Exploring Enhanced Agentic LLM Framework Towards Brain Network Analysis
- Where FactsGo Missing: A LayerwiseTaxonomy and Per-Layer Attribution of Information Omissionin Air-Gapped LLM Agent Pipelines
- TRACE-ROUTER: Task-Consistent and Adaptive Online Routing for Agentic AI
- My Ollama box picks the music now: an agentic DJ running on a 9B model
- Evaluating Agents Beyond the First Prompt
- Reflections and Recommendations on AI Adoption Practice from a Mixed-Ability Research Group
- The Help Ladder: Skill-Adaptive Peer Scaffolding for Real-Time Collaborative Programming
- From Vibe to Code -- and Back: Lexical Oscillation in the Formation of Design Intent with Generative AI
- Beyond Conversations: Spatially-Anchored Previews for Intent Disambiguation in LLM-Assisted Geometry Editing in Virtual Reality
- A Taxonomy of Confabulations and the Perception-Reality Gap in LLM-Assisted Immersive Scene Editing
- Spectral Dynamics of Semantic Drift in Clinical Multi-Agent Language Model Networks
- A Comparative Study of MCP and A2A for Inter-Agent Coordination in LLM-Based Systems
- A Vocabulary for Multi-Agent Automated Research Systems
- How we built LangChain’s agent-first data stack
- The Orchestrator's Tax
- Compliance-First AI: Proving Agent Provenance for Regulated Engineering Teams
- Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI
- Show HN: Formally verified 3D CSG: Trust 93 lines spec, not 1000 lines AI code
- How building software is changing at Anthropic — cited by 1
- Show HN: Tines 3B – safe workflow automation for when everyone builds software
- Scientific computing in the age of agentic AI
- Stop Writing for Me: Generative Refusal in AI Tools for Thought
- Language as a Material Interface for Creative LLM Interaction
- What Gets Lost When Memory Becomes Media? Evaluating AI-Generated Oral History Visualization
- Agentic AI-enabled discovery across large-scale sleep physiology
- SafeFlow: Semantic Information-Flow Control for Blocking Malicious Propagation in Multi-Agent Systems
- ARCHER: Agentic Rule and Compliance Harness for Executable Regulations
- CHILL-Harness: Counterfactual Harness Learning for Efficient Reasoning in Long-Horizon Agents
- Towards a Systems Foundation for Agentic Cloud Management
- Aethel: A Reproducible Graph-Retrieval Framework for Multi-Hop Financial Diligence
- Who Cares About the Model?
- Formal methods with Hillel Wayne
- One Run Is Not an Idea: The Implementation Lottery in Automated Research
- Living-Harness Is an Interactive-Agent Evolver
- DREvo: Distilling Recalibrated Historical Experience for Harness Self-Evolution
- When Should AI Follow? Task Structure and Joint Adaptation by Human and AI Agents
- Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation
- Ontologies Are So Back: Why AI Agents Are Reviving the Semantic Web
- Loop engineer practice #1: Reddit loop grew 0 to 95 Karma in 7 days
- The Economic Benefit of Refactoring
- Stacked pull requests are now in public preview
- Auditing Emergent LLM-Agent Collaboration through Cooperation-Obligation Coupling
- Scaling LLM-Driven Multi-Agent Systems: Design Principles and Architectural Scalability Analysis
- AgentRadio: Passive Awareness for Long-Horizon Multi-Agent Collaboration
- Autonomous Event-Driven Multi-Agent Orchestration for Enterprise AI at Scale
- A Systems Engineering Framework for Vision-Language-Enabled UAV Triage and Disaster Response
- LabEvolver: Training-Free Experience Evolution for Safe and Grounded Wet-Lab Agents
- Univé builds an AI-ready workforce
- Show HN: What should the GUI for AI agents look like?
- Inkling-Small 🧠, GPT-5.6 price cuts 💸, Gemini Robotics 2 🤖
- The Conductor Developer
- Evaluating code review agents with ReviewBench
- Are you ready for Le Chaton FAT or still wasting money on GPUs?
- Claude for ADHD: The Coding Workflow I Built for My Brain
- Unanticipated Effects of Generative AI on Expertise Pathways and Performance Perception in System Administration
- TransMem: Transforming Hidden States into Memory for Large Language Models
- SeekBrain: An Autonomous Multi-Agent System for Accelerating Neuroscience Discovery
- Beyond Byzantine: An Organizational Consensus Algorithm for Self-Interested Agents Under Information Asymmetry
- Autonomous Repair for Multi-Agent Systems via Monte-Carlo Tree Search
- My Super Simple Software Factory (For Agentic Engineers)
- How Stripe Built their Knowledge AI Platform: A Company-Wide AI Agent on Deep Agents, Live in 1 Week
- Show HN: "Hedgehog" - An Opinionated workflow for building with AI
- ReVoicer: Conversational Voice Annotation for Human-Centered, LLM-Assisted Peer Review
- Revibing Code from Papers: Reimplementing HCI Artifacts
- Width, Memory, and Delay: A Resource Accounting for the Limits of Flat Multi-Agent Systems
- MAPLE-Guard: Memory-Aware Link Enforcement Against Memory-Link Poisoning in Multi-Agent Systems
- BANDMAS: Causality-Inspired Semantic Packet Scheduling for Bandwidth-Efficient Multi-Agent Collaboration
- HIERA: Hierarchical Multi-Agent Relevance Assessment for Content Discovery Systems
- Asking Questions the Right Way: A Multi-Agent Conversational System for Prompt Formulation in Complex Task Resolution
- Training Small LLMs as Spatial Multi-Agent Policies
- Customize the reasoning level for Copilot cloud agent
- Trigger Copilot automations with comments
- Customer Experience (CX) Agents in Production: Lessons from Lyft, Vodafone, and LATAM Airlines
- Show HN: Gigacode – the model writes its own multi-agent workflow, then runs it
- How to evaluate voice agents: execution, outcomes, and experience
- Unpacking ChatGPT Work: the Agent for a Billion Users
- Stateful Governance for Concurrent Agentic Systems
- Emergence of Biased Consensus in Multi-Agent LLM Debates
- SABRE: A Multi-Agent Approach for Selecting Out-of-Distribution Detectors Under a Budget
- Group Perspective Matters: Regulating Debate Relationships Can Mitigate Blind Conformity in Multi-Agent Debate
- An Actionable Diagnosis of Multilingual, Multi-Agent Planning Failures
- Agent memory layers don't need an LLM deciding what to remember
- How we build an autonomous SRE Agent for Kubernetes Deployments
- Enacting Constructive Conflicts with AI Agents to Enhance Reconsideration among Novice Interaction Designers
- LEGOUI: Designing with UI-DSL Bricks to Balance Transparency and Controllability
- IntentLint: Supporting Intent Scaffolding and Prompt-time Linting in Human-AI Collaborative Data Analysis
- Preference-Driven Online Adaptation for Personalized Interaction Initiation in Proactive AI Assistants
- Strategic Evaluation of Planning Strategies for LLM Agents in Cyber-Physical Systems
- Continuous Improvement and Parallel Autonomous Exploration: An LLM-Agent Framework for Searching Large Solution Spaces
- HELENA:Hierarchical Sparse Coordination over a Union of Complementary Topologies for MAS
- CURATE: Leveraging LLM Agents to Compose, Catalog, and Deploy Reproducible Workflows
- MIDAS: Multi-LLM Iterative Data-Adaptive Summarization
- Structured LLM Reasoning for Zero-Shot Human--Robot Coordination Under Hidden Goals
- Models, Harnesses, and Multi-Agent Systems
- Crazy AI workflow for banks/finserv – internal audit
- A Two-Tier Perspective on Inference-Time Parallelism in Multi-Agent LLM Systems
- Certifying Collective Reasoning in Multi-Agent Systems via Koopman Spectral Analysis
- ASGE-RR: Agentic Service Graph Embedding with Revisable Reservations for Dynamic AI-Agent Calls
- DoctorAgents: an agentic framework to iteratively refine AutoML pipeline for small clinical temporal data
- Adaptive Arena-based Contestable Argumentative Network-of-Experts for Open-Ended Care Plan Coordination
- How HSP GRUPPE builds AI capabilities for tax advisory
- Your AI Second Brain Is Slowly Rotting (Here's How to Fix It)
- [AINews] Zawinski's Law of MultiAgents
- Speculative decoding in a tools call
- Evaluating XAI Support From A Hierarchical Reinforcement Learning Policy in Human-Agent Collaboration
- Fact-Check Your Information (FYI): A Design Probe to Understand How People Actually Fact-Check Data-Driven Articles
- Strategy-first synthesis planning for complex natural products
- ADIAS: Automated Design of Interactive Agentic Systems
- Reflex: Demonstrate a GUI workflow once, replay it with zero LLM calls
- Model ML completes finance work more efficiently with GPT-5.6 Sol
- What building an AI-native finance function taught me
- The death of AI workflow builders
- How Zapier transformed core marketing processes with ChatGPT Work
- Virgin Atlantic sharpens customer journeys with ChatGPT Work
- From Human-Centered Design to Human-AI Collaboration: Why the Future of HCI Still Starts With People
- MasDrift: Benchmarking Authorization Preservation Across Multi-Agent Architectures
- Fluid Structure, Rigid Record: A Layered Organizational Design Framework for Agent-Native Organizations
- Muscle Memory for Agents: Compile not Merely Retrieve
- Beyond Tier Labels: Role- and Deployment-Dependent Model Substitution in Multi-Call LLM Workflows
- MoRSE: Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts
- You Don't Need To Stay in The Loop: An Agentic Robotics Loop for Robot-Policy Improvement
- TDD inside the agent loop - theater or actual value?
- How many of your agent's calls actually need a frontier model?
- Thinking of ACE? We Can Do It with Fewer Tokens
- Show HN: Synapse – a monitoring SaaS shipped with my Claude Code workflow
- Building monday.com Sidekick: why capable agents need more than just tools
- Copilot memory and Ollama in GitHub Copilot for JetBrains
- Beyond Cash Flows: A Multi-Agent AI Framework for Valuing Clinical-Stage, Cross-Border Biotechnology
- ASCon: A Direction-Aware Reciprocal Agent--Step Contextualization Model for Failure Attribution in Multi-Agent Systems
- Reifying Research Logic: AI-Assisted Workflow Construction and Incremental Refinement for Quantitative Syntax
- Who Are You Explaining To? A Multi-Agent System for Audience-Aware XAI Narratives
- Persistent Recursive Worlds Enable Autonomous Software Evolution
- What is an AI agent?
- Stop being skeptical about AI for development with Charity Majors
- How RingCentral builds AI-native work from engineering to ops
- Socioduality: A Relational Process Framework for Human-AI Interaction
- When Do Institutions Beat Intelligence?
- Rethinking Agent Security as a Networking Problem
- Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop
- MaSRead: Content-Addressed Reading of Replicated Latent Stores
- Harnessing agent memory to build lifelong AI partners for materials scientists
- Conformity Mitigations in Large Language Models Lie on a Single Resistance-Receptivity Frontier
- EvoGraph-Mem: Failure-Aware Editable Graph Memory for Long-Term Language Agents
- Record, train, and deploy from one place with Strands Agents, LeRobot, and Hugging Face Storage Buckets
- What We Learned by Reproducing 2,200 papers from ICML
- Humans are Missing from AI Coding Agent Research
- Interaction Readiness: A Framework for Building and Evaluating AI Agents in Human Roles
- Discovering Efficient and Explainable Communication Topologies for LLM-based Multi-Agent Systems via Causal Inference
- Reconcile Once, Write Anytime: A Trust-Tiered Librarian and a Multi-Agent Writer for Drift-Free, Point-in-Time Research
- Attune: A Self-Annotation Tool for Understanding Robot Operator Attention Profiles
- How to Build the Most Powerful System for AI Coding (Full Breakdown)
- React for Agents: Astro Creator Brings Hooks to his Meta-Harness, Flue
- Proxy-Validated LLM UX Micro-Simulations: An Artifact-First Protocol for Early-Stage Decision Support
- MobileMem: Learning from a Year of Mobile Experiences
- From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL
- A Graph-Based Reinforcement Learning Framework for Structured Drift Diagnosis and Recovery in Autonomous LLM Agents
- The Ultimate Guide to Making Your Entire Development Cycle AI Native
- FIXING Opus 5: PROOF that Prompt Engineering IS NOT DEAD
- AI Agents and the Future of VIS
- RaivenTracks: Branching Provenance for Conversational Visualization Workflows
- Who's Keeping Score? Interactive Steering of LLM-Powered Scoring with Attune
- BRA-Audit: Budgeted Runtime Auditing for LLM Multi-Agent Systems via Cumulative-Exposure Audit-Point Placement
- Emergent Misaligned Communication in Long-Horizon Multi-Agent LLM Commerce
- ETHOS: Towards a Modular Ethics Framework for Clinical Multi-Agent Systems
- VCE-Skill: Enhancing Skill Self-Evolution with Version-Change Experience
- The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipelines
- From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems
- A practical workflow for LLM-assisted development
- How Much Memory Does Your Agent Actually Need?
- Frontier Model Cost and Open-Weights Popularity is Driving Demand for Model Routing
- How NVIDIA scales expertise with ChatGPT Work
- Why This and Not That? A Collaborative Reflection Approach for Understanding Thought Coverage in Decision Making Support Dialog
- Appearing Legitimate is Not Enough: Interrogating Synthetic Agents in Representational Processes through a Participatory Design Lens
- Procedural Collapse: A Structural Account of Disengagement in LLM-Assisted Writing
- MITRE-SAGE: A Multi-Agent Cybersecurity Question-Answering model
- The Little Scientist: LLM Agent-Driven Discovery via the Scientific Method
- KernelArc: A Multi-Agent Framework for GPU Kernel Optimization
- Am I doing something wrong? Qwen 3.8 27B seems useless for agentic coding
- From Chrome DevTools to AI Engineering, with Addy Osmani
- Show HN: Grove, a formal workflow protocol for long-running AI coding agents
- Citizens Build, Agents Execute, Experts Govern
- LEDGER: Claim-to-Evidence Trace Graphs for Auditing LLM Agents
- Measuring Proof Burden in Public Bounty Listings: A RentAHuman Case Study
- Report on The 1st Workshop on Human-Centered Proactive and Personalized Agents for Interactive Information Access at CHIIR 2026
- Contracting for LLM Delegation: Moral Hazard in Technology and Effort Choice
- A Locally Deployable Tool-Grounded LLM Multi-agent Framework for Automating Methane Emission Analysis and Reporting
- Bayesian Partner Modelling enables Adaptive Replanning for LLM Coordination
- Towards Reversible Forgetting: Managing Obsolete Knowledge in Continual Enterprise AI Agents
- CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks
- DentAgent: Evidence-Centric Multi-Agent Coordination for Multimodal Dental Reasoning
- AME Agent Swarms Rewrite the Workflow
- The Pulse: We need to talk about migrations with AI
- AQuA's "self-improvement" updates research state, not the agent LM. What should a local port freeze?
- The /wayfinder Skill: Navigating the “Fog of War” of Planning
- I did it! I'm free! It's been 7 hours since I used claudecode
- Delegating or Doing? Understanding User Behavior in Hybrid Human-Agent Interfaces
- IRIS: Navigating and Reflecting on Writing Traces Using Intelligent Document Histories
- Reward-Guided Autoregressive Graph Generation for Efficient Multi-Agent Communication Topology Design
- Hallucination as a Feature, not a Defect: Evaluating a multi-agent architecture to transform speculative language-model outputs into testable scientific hypotheses
- When Do LLM Agents Help? Deadline-Aware Mixed-Criticality Task Scheduling at the Autonomous-Vehicle Edge
- Show HN: A coding-agent workbench built around Matt Pocock's coding workflow
- Recursive Self-Improvement
- The new GitHub Copilot experience in Slack
- I tried to do agenic coding with Qwen 3.8 27B 3bit quant on a macbook air m2 24gb. It took 63 hours, but amazingly, the flight simulator worked.
- Show HN: I let an subagent workflow refactor my codebase for three days
- The Evolution of the Agent Harness
- Fast and Hard Code
- Adapting Fossil-scm as a platform for AI agentic workflow
- Has anyone actually made 64k feel like 300k+ with recursive local agents?
- 1/100 → 44/100: fine-tuning a 450M VLM on 50K browser screenshots
- Human judgment doesn't leave the software factory. It relocates.
- Practical Loop Engineering
- Agentic Code Quality
- The Belief Update Gate: Separating Inertia from Learning in Human-AI Interaction
- Edge-Based Agentic Retrieval-Augmented Generation for Autonomous FHWA Bridge Inspection Compliance
- Towards Traffic Modelling of Multi-Agent Systems: The Role of Coordination Topology
- Complete Cyclic Subtask Graphs for Tool-Using LLM Agents: Flexibility, Cost, and Bottlenecks in Long-Horizon Workflows
- Intelligence EXPLOSION: Harness Engineering with Pi Agent, Deepseek, and Gemini
- Fragments: August 24
- How Toyota North America Put Enterprise AI on the Balance Sheet with Deep Agents and LangSmith
- Context Anchoring
- Humans and Agents in Software Engineering Loops
- Ask HN: Is there any way to use workflow to control Harness?
- Exploring Agentic Approaches for Data Issue Detection and Repair in AI-Assisted Visualization
- Multi-Agent Discovery and Resource-Aware Autonomous Exploration of Scientific Datasets
- Probing How Users Interact with Turn-Level Design Frictions for AI Chatbots
- "I want to be pushed, I want to grow": Enabling social workers to design evaluations of LLM augmentation in their work
- Opinion-Guided Layered Strategies for Decentralized Coordination
- PropUQ-MAS: Propagation-Aware Uncertainty Quantification for LLM Multi-Agent Systems
- Right-Sizing LLM-Agent Decomposition in VAT Determination: A Pilot Controlled Sweep
- The Interaction Tax: When Communication Erases Diversity in Multi-Agent Teams
- Agentic Scaffolding Amplifies Sycophantic Behavior in Large Language Models
- Today I merged the first feature branch written entirely by my 4060Ti 16GB!
- AI Proficiency: From Users to Builders
- How We Build Agent Environments & Tasks
- Building Self-Correcting Memory in OpenWiki
- Why Ramp built its own in-house coding agent, Inspect
- Markets, Not Planners: Decentralized Orchestration of LLM Agents with Private Information
- Agentopia on a Consumer GPU: A Reduced-Scale Long-Horizon Port with an 8B Model
- LLM Agents Perform Controlled Experiments Using Simulation Models
- Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment
- MARS: Multi-Specialist LLM Relay System for Competitive Programming
- How loveholidays is making everyone a builder with Codex
- A minecraft clone I fully vibecoded with Qwen3.8-27b Q4
- Watch This If Your Coding Agent is Ignoring Your Rules (You Need Hooks)
- Agentic World Analysis (AWA) - an alternative way to explore systems and support decision making
- HypoForge: A Self-Improving Multi-Agent Framework for Automated Hypothesis Generation and Testing via Scientific Skill Learning
- Praxist: From Experimental Artifacts to Solution Lineages
- Poisoning Agentic Alpha: Adversarial Vulnerabilities Across Roles and Architectures in Multi-Agent Trading Systems
- Federation Is Nearly Free, Reasoning Is Not: Tradeoffs for AI Co-Scientists in Protein Characterization Workflows
- MACGen: Toward Functionally Correct and Secure Code Generation via Multi-Agent Collaboration
- Beyond Scaling: Self-Evolving LLM Agents for Hardware Kernel Optimization via an Experience-Driven Workflow and Experience Graph Memory
- My software development workflow is AI now & it feels exhausting and soulless
- Making Your Data Ready for Agentic AI
- Show HN: Build your own theme park
- Improving LLM Interpretability with User-Centric Chain-of-Thought Reasoning
- Zero-Shot Self-Orchestration with Ledger-Based Control for Improved LLM Coding Performance
- Risks and Controls for Multi-Agent Systems: an analytical framework for deployment of AI agents across organisational boundaries
- One Model, Many Minds: Unlocking Multi-Agent Synergy in a Single Agent via Mixture of Roles
- Agent Mesh: Reliability Primitives for Non-Idempotent Agent Delegation - Identity Adequacy and Evidence Adequacy
- SKILL.state: Scalable Long-Horizon Agent Skills
- MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models
- Building the Foundation for the Agentic AI Era
- Best workflow engine is a programming language
- Agentic Workflow Design: Six Principles for 2026
- AI Can't Replace Real Research in Empathy Mapping
- The Custodial Era of UX: Cleaning Up After AI
- The Hidden Flaw of EVERY Coding Agent Now Has a Solution
- Agency and Agents
- How Much Can AI Understand? Toward AI-Assisted Sensemaking of Collaborative Discussion in Groups with Shared History
- FocusGen: Expanding Visual Design Exploration with a Simulated Focus Group of Persona Agents
- AI as Teammate: Rethinking Task Distribution in Medical Training
- Between Algorithm (AI) and Intuition (Human): Preserving Designer Agency in AI-Assisted Sensemaking of Qualitative UX Data
- FedEHR-Agents: Federated Agentic Optimization for Automated EHR Modeling
- Prove2Me: An Open Collaborative Platform for Scaling Math Formalization
- Offline-Verifiable Accountability for Cross-Organization Agent Messaging: A Preserved Evidence-Bundle Approach
- Logos: An Agent Harness on a Cross-Process Bus
- Space to Talk
- Agentic Engineering Operating Level: WHERE to FOCUS your AGENTS?
- GLM 5.3 and GLM 5.3 Flash ran locally on RTX PRO 6000 WS and built a penthouse using BlenderMCP
- Delegating Before Learning: Where Generative AI Sits in Students' Professional Communication
- Structured State Reconciliation for Human-AI Task Handover
- AREAs-Lab: An Interactive Environment for AI-driven Requirement Elicitation for AI Systems
- ASTRA - Agentic System for Ticket Resolution and Analysis
- Forward-Deployed Full-Stack Engineering for Autonomous Cloud MLOps
- AgenticRag-R1: Agentic Reinforcement Learning with Stack Memory for Multi-Step Reasoning, Retrieval and Memorizing
- Harness-RL: Black-Box Reinforcement Learning with Action-Args Decoupling for Central-Agent Multi-Agent Harnesses
- Cognitive Cells: A Compositional Framework for Populations of Small Language Models
- How Foundational Models Became Superhuman in Bash
- PRs NOT Welcome: How Top AI Open Source Projects Are Managing Thousands of Contributors
- How AI-native companies turn workflows into operating capability
- Fragments: September 1
- Copilot code review can now approve pull requests
- Ask HN: What full workflow / process have you tried to automate with AI?
- How law firm Gilbert + Tobin governs and scales AI with OpenAI
- Are We There Yet? Assessing Computer-Use Agents for Blind Users' Accessible Interaction with Desktop Applications
- Classic AI Scaffolding for LLM Social Agents
- Long-Horizon State Tracking in LLMs: Executing MD5 through a Deep Sequence of Dependent Tool Calls
- EULER: Exploring Underused Links with Evidence-Checked Return for Multi-Agent Mathematical Discovery
- Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs
- Qwen 3.8 Flash Next + HERMES AGENT = AWESOME LOCAL AI AGENTS!
Agent tools & setup
- anti-patterns and patterns for achieving secure generation of code via AI
- Note #727
- Claude Code and What Comes Next
- Configurable AI Coding Assistants: Designing For Developers Who Like to Be in Control
- Introducing Zapp
- An update on recent Claude Code quality reports
- Quantifying infrastructure noise in agentic coding evals
- Beyond permission prompts: making Claude Code more secure and autonomous
- Introducing advanced tool use on the Claude Developer Platform
- Writing effective tools for agents — with agents
- Desktop Extensions: One-click MCP server installation for Claude Desktop
- Equipping agents for the real world with Agent Skills
- Introducing OpenWiki Brains, general-purpose wiki memory for agents
- Introducing Contextual Retrieval
- How Schneider Electric Built Their LLMOps Foundations At Enterprise Scale With LangSmith
- LangChain and NVIDIA launch the NemoClaw Deep Agents Blueprint
- Deep Agents Code on NemoClaw: a governed blueprint for your most sensitive code
- Harbor x LangChain: A Unified Stack for Evaluating Agents
- Introducing OpenWiki, an open source agent for repo documentation
- How Pendo used LangSmith to trace Novus from user behavior to code fixes
- Running Untrusted Agent Code Without a Sandbox
- Prompt Caching with Deep Agents
- Raising the bar on SWE-bench Verified with Claude 3.5 Sonnet
- Rebooting Enterprise AI with MCP and Kubernetes
- Hermes Agent: Agents that grow with you
- Technical advances in document understanding
- Chris on AI, autonomous swarming, home automation and Rust!
- Beyond note-taking with Fireflies
- Show HN: Orchestrator – a single-binary workflow orchestration tool
- Ask HN: Would filesystem bookmarks be useful in your shell workflow?
- Bytechef open source platform for AI agent orchestration and workflow automation
- Show HN: DonnyClaude – a verified workflow engine for Claude Code
- Workflow State Engine – Ferricstore
- Show HN: Wayflow – an embeddable AI workflow builder (open source)
- Chainything: Workflow automation tool with no-code UI and AI assistant
- AI Specialists Ready to Transform Your Workflow
- Open-source AI agent workflow for auditing Solidity smart contracts
- Show HN: What if your menu bar was a keyboard-controlled command center?
- Your coding agents are a black box. Here's how to crack them open.
- Agents need their own computer. Here's how to give them one safely.
- New in LangSmith Fleet: Bring agents into Slack in one click
- AgentSociety 2: An Integrated Research Environment for Executable Social Science
- Ontology-Amplified Distillation and Contextuality Auditing for Sovereign Enterprise Language Models: A Combined Proof-of-Mechanism and Negative-Results Method Study
- SoftBoard: A Multi-Agent Tool for the Creation and Evaluation of Low-Fidelity Prototypes
- DevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments
- OpenWiki 0.2 is adopting the OKF support
- Unigent SDK – universal cross-harness, cross-session agent workflow scripting
- Show HN: Skillful, stop maintaining the same AI workflow in five places
- Show HN: AI Workflow Builder App Template for React
- Show HN: How you auto recover your Claude Code workflow when quota resumes
- Creating a private AI assistant in Thunderbird
- ChatGPT is now a partner for your most ambitious work
- Samsung Electronics brings ChatGPT and Codex to employees
- New usage analytics and updated spend controls for enterprises
- Show HN: Leaves – A text-UI disk usage treemap visualizer
- Show HN: Nobie – an Excel-compatible runtime for agents and humans
- Show HN: Clawk – Give coding agents a disposable Linux VM, not your laptop
- Show HN: Jacquard, a programming language for AI-written, human-reviewed code
- Show HN: Juggler – an open-source GUI coding agent, by the creator of JUCE
- The Pulse: Grok’s CLI caught uploading all your local files to the cloud
- Building OpenCode with Dax Raad
- Beta: Clawk
- Beta: dcg
- Beta: Mellea
- Thinking Machines Inkling 🧠, GPT-Red 🔒, Perplexity sandboxes 🛡️
- Claude Code browser 🌍, Cursor general agent 🤖, Claude Fable extension ⏳
- Devin Fusion 💻, DeepSeek DSpark ⚡, economy of tokens 💰
- Seedance 2.5 🎥, guide to Fable ✨, OpenAI preps GPT-5.6 🚀
- Profiling in PyTorch (Part 3): Attention is all you profile
- From Hugging Face to Amazon SageMaker Studio in one click
- Run AI workloads on any cloud, store on Hugging Face: zero-egress storage with SkyPilot
- 🤗 Kernels: Major Updates
- Run a vLLM Server on HF Jobs in One Command
- Experimenting with the proposed Cross-Origin Storage API in Transformers.js
- We got local models to triage the OpenClaw repo for FREE!*
- From the Hugging Face Hub to robot hardware with Strands Agents and LeRobot
- Agentic Resource Discovery: Let agents search
- Migrating Your GitHub CI to Hugging Face Jobs
- The Open Source Community is backing OpenEnv for Agentic RL
- Designing the hf CLI as an agent-optimized way to work with the Hub
- wtf is Loop Engineer & how to setup for real
- New AI coding paradiagm - OpenAI Symphony
- Anthropic killed Tool calling
- WebMCP - Why is awesome & How to use it
- How to install and use Claude Code Agent Teams (Reverse-engineered)
- Your OS Changes Everything for Local AI
- This Is What Happens When You CRUSH An AI Video Model
- Visualization Autocomplete: Visualization Authoring via Stepwise Design Recommendations
- SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents
- A Generative Partially Specified Finite State Machine Approach to Complex Behaviour Planning
- IssueBench - How We Evaluate Engine
- Octo-planner: On-device Language Model for Planner-Action Agents
- ETAS: An Effect-Typed Language for Agent Systems
- Self-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?
- PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution
- Grabette: an open system to record robot-manipulation data
- The PERFECT Local AI Setup
- Hermes Agent the BEST Local Ai Agent
- Can YOU Run Deepseek V4 Locally?
- Viability of local models for coding
- Evals Skills for Coding Agents
- Selecting The Right AI Evals Tool
- Engineers... STOP Picking GPT-5.6 Sol OR Claude Fable 5… FUSE THEM
- SEE CMUX SOLVE Multi-Agent Orchestration (Claude Code and Pi Agent)
- PLANS For Fable 5: Rebuilding My /Plan Skill for Mythos Class Models
- Engineers, DELETE the BASH Tool: Agentic Security For Pi Agent and Claude Code
- My M5 Max, Gemma 4, MLX LOCAL Stack. (This KILLS MODEL PROVIDERS)
- Building Managed Agents That Use GitHub Without Exposing Your Token
- Control an Android Phone with Gemini 3.5 Flash Computer Use
- Getting started with the Gemini Interactions API
- How Gemini Managed Agents Works under the Hood
- Gemini Managed Agents: Developer Guide
- How to use Deep Research with the Gemini API
- How to correctly use MCP servers with your AI Agents
- 8 Tips for Writing Agent Skills
- Combine Built-in Tools and Function Calling in the Gemini Interactions API
- Practical Guide to Evaluating and Testing Agent Skills
- Writing a Good AGENTS.md
- Multimodal Function Calling with Gemini 3 and Interactions API
- Getting Started with Gemini Deep Research API
- Gemini Interactions API Quick Start
- MCP is Not the Problem, It's your Server: Best Practices for Building MCP Servers
- Building Agents with the Gemini Interactions API
- Introducing MCP CLI: A way to call MCP Servers Efficiently
- Practical Guide on how to build an Agent from scratch with Gemini 3
- Gemini API File Search: A Web Developer Tutorial
- Build your first AI Agent with Gemini, n8n and Google Cloud Run
- This Completely Changes the Way We Build Production AI Agents (Vercel Eve)
- I Turned Claude Code Into a Complete Video Generation System (with Archon)
- My AI Memory Now Follows Me Across Every Tool!
- Agentic Batch Changes is now in public beta
- Why your migration tools are failing your engineers
- Sourcegraph MCP server and a cheaper model beat a Mythos-class model alone
- Lessons on UX, security, and scale when building an enterprise-grade Slack agent
- Code Search, Deep Search, or MCP: When to Use Each
- MCP stories from the field
- A new era for Sourcegraph: The intelligence layer for AI coding agents and developers
- How our support engineers use Deep Search to investigate customer issues faster
- Why code search at scale is essential when you grow beyond one repository
- Fixing the React2Shell vulnerability in large and complex enterprise codebases (part 2)
- Omnigent: The New Meta-Harness for EVERY Coding Agent - Claude Code, Codex, Pi, More
- Google's Agents CLI: The CLI + Skills Combination to Ship AI Agents EASILY
- Better Models: Worse Tools
- Pushing Local Models With Focus And Polish
- Beta SDKs for the 2026-07-28 MCP Spec Release Candidate Are Here
- Enterprise-Managed Authorization: Zero-touch OAuth for MCP
- The 2026-07-28 MCP Specification Release Candidate
- Tool Annotations as Risk Vocabulary: What Hints Can and Can't Do
- Understanding MCP Extensions
- The 2026 MCP Roadmap
- MCP Apps - Bringing UI Capabilities To MCP Clients
- Exploring the Future of MCP Transports
- MCP joins the Agentic AI Foundation
- One Year of MCP: November 2025 Spec Release
- MCP Apps: Extending servers with interactive user interfaces
- Adopting the MCP Bundle format (.mcpb) for portable local servers
- Server Instructions: Giving LLMs a user manual for your server
- Update on the Next MCP Protocol Release
- Introducing the MCP Registry
- Announcing the Official PHP SDK for MCP
- Meet Puck
- Amp Is Now In Slack — cited by 1
- Subscriptions, At Last
- From Agent to Agent
- Secrets of the Orb
- The Dial
- Agents, Anywhere
- More Orb Sizes
- Read Bigger Threads
- Agents in Orbs
- Custom Agents
- A Faster Librarian
- Diffs
- Faster Deep & Rush
- Agents, Everywhere
- Opus 4.8
- The End of Public Threads
- Plugins, Everywhere
- Drop the Neo
- Proof of Human
- GPT Image 2 Paints Better
- Rush, 2.0
- npm Package Changes
- Amp, Rebuilt
- GPT-5.5 In Deep
- Opus 4.7
- GPT‐5.4 in Deep
- GPT-5.4, The New Oracle
- GPT‐5.3‐Codex
- Liberating Code Review
- Slashing Custom Commands
- Painter
- Hidden Gems: Part 4
- What GitHub Copilot's Usage-Based Billing Means for Zed Users
- Terminal Threads Are Live in Zed
- Why and How to Run Local Models in Zed
- Use Your ChatGPT Subscription in Zed
- What Anthropic's New Claude Billing Means for Zed Users
- Introducing Zed for Business
- We're Not Building AI Features for the Money
- Introducing Parallel Agents in Zed
- How We Developed Zeta2
- We Rebuilt Zeta from the Training Data Up
- Choose Your Edit Prediction Provider
- The ACP Registry is Live
- Run Your Project in a Dev Container, in Zed
- Hidden Gems: Part 2
- Introducing Agent Extensions
- Codex is Live in Zed
- GitHub Code Quality is now generally available
- Repository-level GitHub Copilot usage metrics generally available — cited by 1
- Copilot code review: Customization and configurability improvements
- GitHub Mobile: Fix pull request comments with Copilot cloud agent
- Trace voice agents in LangSmith
- Gemini 3.6 Flash is now available in GitHub Copilot
- pi 0.81.0 adds support for llama.cpp
- Show HN: SciStudio: organize chaos data analysis pipelines with workflow runtime
- Right on Schedule
- Copilot users can now see AI credits used per billing cycle
- I ran Laguna-S-2.1 through my private agentic eval vs Qwen3.5-122B on an RTX Pro 6000 (96GB). Fastest 100B+ I've tested and the best tool calling, but it invents facts under pressure.
- AI Tool Discovery at Scale: All You Need is DNS
- Show HN: CodeAlmanac – Karpathy-style codebase wiki from your conversations
- Unsloth Quantization of Laguna S 2.1 Is Out
- Gigatoken: A new open source tokenizer ~100x faster than Tiktoken, -500-1000x faster than Huggingface
- Introducing OpenAI Presence
- Llama.cpp just added support for Laguna XS.2 & M.1
- Towards Automating Eval Engineering
- Show HN: Ipek – a visual IDE for workflow automations
- Show HN: Bento - An entire PowerPoint in one HTML file (edit+view+data+collab)
- New Copilot usage metrics impact dashboard
- MindControl - llama.cpp fork to guide the reasoning process via injection during sampling
- Tool: Databasement
- Beta: CodeAlmanac
- Beta: termcn
- Show HN: Cactus Hybrid: We taught Gemma 4 to know when it's wrong
- How We Benchmark Deep Agents
- Code Finder: fast, efficient code search for coding agents
- Event Driven Orbs
- PSA on Laguna S-2.1 - Use the updated chat template and GGUF
- July 2026: LangChain Newsletter
- Show HN: OneCLI – OSS credential gateway that keeps secrets out of AI agents
- Show HN: Palmier Pro – Open-source macOS video editor built for AI
- Show HN: Remux – an open-source tmux workspace designed for iPhone
- Show HN: DeepSQL – A self-hostable DBA agent for Postgres and MySQL
- GitHub Mobile: Fix failing Actions checks with Copilot cloud agent — cited by 1
- Agent automation controls in GitHub Issues in public preview — cited by 1
- GitHub MCP Server supports the next MCP specification
- How Codex became a collaborator for OpenAI’s creative team
- I compared local models and different quants / config on a subset of swe-verified bench
- [audio.cpp] Release 0.4: Higgs Audio v3 TTS 4B (10x real time)+ Fish Audio S2 Pro in C++/GGML, full GGUF loading, Q8 speed and VRAM gains
- HARP: The Human--AI Research Platform
- AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use
- UPDATE - HuggingHack Is Now On Github
- Using the Bonsai 27b 1b quant locally - regularly.
- Demios – AI-native workspace for data capture, workflow automation and reporting
- CachyLLama’s: llama.cpp fork with persistent KV cache that makes long local-agent sessions much less painful
- DKV: Open-source KV-cache compression framework for local LLM inference (CLI + technical report)
- Python Toolkit: a GUI to manage python, venv, packages, reqs, AI interfaces and more...
- CachyLLama: llama.cpp fork with persistent SSD-backed KV caching for local agent workflows
- Show HN: Claude-thermos keeps your Claude session warm for you
- Llama.cpp now has full MCP support!
- POCKET-35B agentic model on cpu 59 t/s
- Minimax M3 support with MSA has been merged into llama.cpp
- Harness showdown: Claude Code vs OpenCode vs Pi with DeepSeek V4 Flash
- Enterprise managed settings in the GitHub Copilot app and Copilot cloud agent
- Nifer is insane. 700t/s with Qwen 3.6 35B (no thinking). Purpose build for RTX5090. Full 250k context too.
- Decentralized Granular Access Control for Agentic AI Systems in Critical Infrastructure
- CRAFT: Learn the Schema, Execute the Plan
- Real2Sim2Real for Vision-Language-Action Manipulation: An AMD ROCm-Based Pipeline
- GitHub Copilot for JetBrains adds improved OpenTelemetry configuration and model management
- Show HN: Yap – OSS on-device voice dictation for macOS with no model to download
- Show HN: Whetuu – a zero-config cross-shell prompt written in Zig
- Show HN: I left VSCode to build an IDE to handle many projects/agents workflow
- GitHub Actions holds potentially malicious workflows for approval
- spec: add DSpark speculative decoding by wjinxu · Pull Request #25173 · ggml-org/llama.cpp
- Gemini Distillation Service
- DeepSeek V4 Flash, up to 32 tok/s on AMD Ryzen AI MAX+ 395
- The 2026-07-28 Specification
- Grok 4.5 is now available in GitHub Copilot
- I got Kimi-k3 running.....
- The Anthropic Economic Index connector
- Your OS Changes Everything for Local AI
- I tried running a 1.56TB MoE model on a 6GB RTX 4050 Laptop, Here’s the result
- Appliedin: The agentic workflow for applying jobs, so we can spend time prepping
- I built a GBNF grammar compiler that makes 8B models reliably call tools - here's how it works (deep dive)
- AI slowdown pact ⏸️, Personal superintelligence access 🌍, Grok Build Mode 🛠️
- The idea: on a CPU the decode speed depends on the active params per token, not the total. My objective is trying to run a 10B at 100tok/s on a mid level PC (No GPU).
- A slide deck you can edit with a local model or in Chrome — the whole deck is a JSON block in one HTML file (~640KB with editor and viewer included)
- How Similarweb Evaluates Long-Form Agent Research Reports with LangSmith
- From keyword to published post in one focused workflow
- Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
- Show HN: Bullshit Detector – agent skills that fact-check videos and articles
- Deep Agents v0.7
- Everyone posts day-one impressions. What's still in your stack a month later?
- Kimi K3 for local use (1.56TB → 594GB) compressed and released by Unsloth
- PSA: llama.cpp now loads MTP tensors by default for any draft-mtp arch, even with MTP disabled
- Ilintar's Official Guide To Model Selection
- Show HN: Qwen Scribe – local transcription and dictation for Apple Silicon
- Copilot code review: Agent skills and MCP now generally available
- Tool: superfile
- Beta: OpenWiki
- Beta: jcode
- The Ultimate Knowledge Base: Bring YouTube Into Your AI Second Brain
- Bought a 5090 to escape API fees. Ended up building a mini datacenter. Sound familiar?
- Anyone tested the IQ1_M 342GB Pruned Kimi K3? Is it usable?
- Benchmarked: MindControl for Llama.cpp
- Turbo-fieldfare: Open-source engine running Gemma 4 26B in 2 GB RAM on Apple Silicon
- GitHub Copilot in Visual Studio — July update
- LangSmith LLM Gateway: runtime controls for production agents
- Reference same-repository actions with self-repository syntax
- Limit remote control to managed devices
- GitHub Copilot in Visual Studio Code, July 2026 releases
- Show HN: Optimize and serve models with Fable quality at half the cost
- How avatarin built a 24/7 retail agent with GPT-Realtime
- VizPilot: Automated Onboarding for SVG-based Composite Visualizations using Multimodal LLMs
- Argonaut: Interactive Visual Exploration for Distributed Optimization
- VISA: A Structured Description Protocol for Agent-Based Simulation Models Towards Machine Reproducibility
- The Complete Local AI System with A Single NPM Install!
- Enterprise teams model policy targeting in public preview
- Show HN: Claude-account – switch Claude Code accounts without logging in again
- How to evaluate Sourcegraph on your own codebase
- [audio.cpp] Release 0.5: DramaBox expressive TTS, Confucius4 cross-lingual voice transfer, plus 7 more models and ROCm/HIP
- Weight-Aware Streaming Tensor Engine: run Kimi K3 using 29 GB of RAM at 0.50 tok/s
- DeepSeek V4 Flash 0731 IQ2_M benchmark for Dual 3060 and 96GB RAM ≈ 3.5 tok/s.
- Fix for Deep Seek v4 Flash 0731 tool calling has been added to llama cpp
- DeepSeek-V4-Flash-0731 UD-IQ3_S 12.5 tok/s on RTX 3090 +128GB DDR5
- Koboldcpp v1.118 released
- Ran DS V4-Flash-0731 Locally on 3xMI50 32GB @ ~15 t/s TG
- I pushed Kimi K3 onto one CPU with 8 GB of RAM
- DeepSeek-V4-Flash 284B on 5.3GB of memory
- Setting up of a 16xGB10 (DGX Spark) cluster
- llama.cpp just added MTP / DSpark support for DeepSeek V4 Flash
- Deepseek-V4-Flash-0731 Dwarfstar on Mac
- Show HN: Sprocket – The Best AI Agent for Hardware and Software Development
- Show HN: NixOS-DGX-Spark – Nix and NixOS on the DGX Spark
- CyberNeuro: A Privacy-Preserving Agentic Workbench for Cohort-Scale Neuroimage and Clinical Data Analysis
- The AnyLog Edge Data Fabric
- Show HN: Mu – Tools for Agents
- I CANNOT believe I've got DeepSeek-V4-Flash-0731, a frontier model, running on my home PC. Insane!
- DeepSeek V4-Flash (284B MoE) at 33 tok/s single / 68 tok/s aggregate on 2× RTX 3090 + a used quad-Xeon DDR4 server — full config
- Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
- can someone create a website where people share specific hardware specs with specific llama cpp flags so we see what works?
- Attach Anything
- Time to finally migrate from LM Studio -> llama.cpp, your experience?
- Sharing a persistent browser QA workflow for OpenCode
- Code coverage automatic enablement in Code Quality settings
- Deploy local agents everywhere with LFM2.5-2.6B
- Show HN: Fine-tune an 8B model on a 4 GB laptop GPU
- [Deepseek-V4-Flash-0731] Full 1M context on a single RTX5090 + DDR5 Desktop Setup with VLLM CPU/Ram Offloading, ~800 tps pp & 15+ tps decode [Agentic Coding]
- Gemma 4 on 500MB
- Has anyone tried Mach-1 Additive? 95% of performance of Qwen 3.6 35B while being 10x smaller
- A llama.cpp PR caches “hot” MoE experts on the GPU — 33 → 56 tok/s reported with 8GB VRAM
- A 2.6B model with tool calling and 128K context now runs at 30 tok/s on a phone
- [AINews] Megakernels are so dead and so back
- MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale
- Qwen3-TTS voice cloning is now in mainline llama.cpp — the old demo finally became real support
- Google LLM router ➡️, Cloudflare Wallets 💳, Anthropic and Volta 🤝
- Sandboxing
- Could we have a --disk-moe or --n-disk-moe like --cpu-moe or --n-cpu-moe so we can use disk/cpu/gpu ?
- Tool: Mu
- Beta: @cloudflare/computer
- Beta: key-amnesia
- The Creator of Claude Code Said to Do What Now?!
- Prime Agent - a new coding harness surpassing Codex/CC/PI
- EDATracer: An Agentic Framework for Large-Scale EDA Artifact Analysis
- Portals into Orbs
- i just spent weeks rewriting my webUI from scratch, getting rid of all AI slop within the codebase and switching it over to a proper lightweight framework (alpine.js). i am now comfortable suggesting it as an alternative to openwebUI, librechat and the like! it is made for local models
- Auto-fit vs tuned MoE offload: 564 → 1330 pp tok/s, unchanged decode (Qwen3.6-35B-A3B Q6 / RTX 3090)
- Best llama cpp flags to run Deepseek-flash 0731
- I compared even more parsers on 14 PDF-parsing capabilities using different types
- nvidia/NVIDIA-Nemotron-Parse-2.0 · Hugging Face
- I ported vLLM's serving stack to C++20: 66 MiB binary, no Python at inference, output checked token-for-token against vLLM
- Deep Agents vs LangChain vs LangGraph
- Kimi K3 is now available in GitHub Copilot
- Show HN: The Channels SDK – Bring Any Agent to Any Channel (Slack, MS Teams)
- 🟩 NVIDIA's whole speech stack just went local. ASR + TTS + codec, quantized to GGUF, running on-device via NeMo-Speech.cpp
- A llama.cpp PR makes Q2_0 3.0–3.6x faster on x86 CPUs, 8B decode goes 2.39 → 8.20 tok/s
- Size the Orbs of Production!
- Show HN: A workflow for building community skill catalogs
- GitHub Code Quality no longer adds Copilot as a reviewer
- Managed Deep Agents is now in Public Beta
- llama.cpp PR reports up to 169% faster quantized-KV decode at 118K context on Intel Battlemage from one SYCL kernel switch
- Copilot usage metrics API adds agent app activity
- MCP allowlists in enterprise managed settings
- Copilot code review effort levels are generally available
- Qwen 3.6 27B flags/settings in llama.cpp
- GitHub Copilot weekly releases — August 3
- Serving Deepseek v4 Flash 0731 on 2x DGX Spark — 5-7 GB OS headroom, what would you do to lower VRAM usage and increase OS available RAM?
- I got tired of my 300GB model loads taking 5min on RPC. PR 26291 speeds it 300% to 1min30sec (4060ti+ddr4) + (4060ti+ddr5)
- Qwen 35B-A3B MoE vs 27B dense in local coding tests: ~4× faster, much smaller quality gap than I expected
- Qwen3.6 27B + 35B on vLLM, single R9700 (gfx1201)
- Claude Code in 9 lines python
- Building a zero-dependency C inference engine for BitNet (1.58-bit) - lessons from hitting 36 tok/s on a Xeon CPU
- Show HN: A terminal glued to the macOS dock
- Kimi K3 (Unsloth) IQ2-XXS from 711GB down to 478GB!!! Only Multi-language was removed to trim the size
- Extremely slow DSpark draft model performance (1-2 t/s) with DeepSeek-V4-Flash on llama-server compared to MTP?
- ds4 flash 0731 UD-IQ2_M wrote a custom metal kernal for kimi k2 IQ1_0 in about 50 minutes
- Best Embedding + Reranking Model
- AMD llama.cpp: reducing MTP buffer overhead gave me 64K → 149K context for Qwen 27B
- DeepSeek v4 Flash 0731 locally on CPU
- Two flags took the official Ling-3.0-flash INT4 from 20.8 to 38.7 tok/s on one DGX Spark
- [NEW MODEL] SupraElegans-500K
- A Dial for You
- Introducing Muse Glimmer: an open-weight model optimized for always-on local agent workflows
- 1M context with 17 GB model in 24 GB VRAM: "for the first time I was able to load a context of almost 1M tokens and extract 7 needles from various parts of the text"
- Engineers… Your Software Factory NEEDS Agent Sandboxes to SCALE (exe.dev)
- OpenAI Astra pause 🚨, Claude Code cross-session 🤖, how Cursor Router works 🔀
- Copilot on web expands conversation controls
- Best Local LLMs - August 2026
- Build Low-Latency Multilingual Voice Agents: Open Weights & Full Deployment Control with NVIDIA Magpie TTS
- DeepSeek V4 Flash 0731 is the ‘killer app’ that is going to sell A LOT of DGX Sparks
- Show HN: Ante, a coding agent in a single binary that runs offline
- I compared GGUF quants of Qwen3.6 27B to NVFP4, AWQ, AutoRound, and FP8
- Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots
- Meta Muse Glimmer 30B Local AI Review
- Global Plugins and Skills
- Show HN: Mcptoon – Token-efficient MCP CLI client
- Introducing Unsloth Desktop app
- You Don't need to use Cloud AI! Switchyard and Nemotron 3.5 Lightning
- How to Dogfood Your AI Chat Agent: A Three-Layer Evaluation Framework with Goal-Directed NPC Simulation
- Automating and Scaling Behavioral Scientific Research on AI Agents
- FYI: Muse Glimmer Chat Template Got Updated Recently
- LangSmith BYOC is now generally available on AWS
- Google Maps and Google Search now work together in the Gemini API
- Introducing Delta
- Agent Plugins 1.0 in VS Code, Copilot CLI, and the Copilot app
- Meta's Muse Glimmer 30B now runs up to ~3.3x faster on Mac with mlx-dspark
- Why managed agents are the next big thing in agent building
- Show HN: Ballet – Workflow automation that writes integrations against any API
- Tool: Amp
- Beta: git-knife
- Beta: Docker Sandboxes
- Every Claude Code Skill I Use to Drive My Entire Development Process
- How do you plan to run Qwen3.8-2.4T-A95B locally?
- Beyond Memory: A Transactional Continuity Kernel for Long-Lived AI Agents
- Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration
- The builder’s guide to GPT‑5.6
- Gemini 3.7 Flash is now available in GitHub Copilot
- Trained a 1.5B to write shell commands so I'd stop googling tar flags. Runs on a laptop CPU in ~1 sec.
- Fixed Jinja chat template for Qwen 3.5, 3.6, and the new 3.8 release
- Show HN: MCP Memory – Fast Agent Memory Using Google's OKF and SQLite FTS5
- Show HN: MCP-stama – An ultra-fast Rust MCP server with no dependencies
- LFM 2.5 2.6B is the best small model for tool use I have ever used.
- bitsandbytes creator teasing new quantization method: GLM 5.3 on a single DGX Spark at 7t/s
- fantastic: latest llama.cpp server webui can now run commands for tools into rootless sandboxed containers
- Senv: Sandboxed Python environments with the uv workflow
- Grok 4.6 is now available in GitHub Copilot
- GitHub Copilot weekly releases — August 10
- Local uncensored Opus 4.6 at home - Qwen3.8 27B heretic
- Show HN: Mole – Deep research agent for your terminal
- Show HN: Ember – Redshift safe color palettes
- Flownie – Open and Visual Data Workflow Platform with AI Agent Assistance
- Fable 5 refuses to touch Qwen deployments?
- Show HN: ThoughtDAG – An editable context graph for LLM conversations
- A nice local vision test
- SOTA Apple Silicon Inference (August 15, 2026)
- Show HN: Deltix – AI Driven Testing
- If you are at the lowest budget, which you can think of.Which hardware would you recommend to run? qwen 3.8 27b oWith like 50 tokens per second. I currently have a RTX 5070 Ti.
- Qwen3.8-27B Hybrid IQ4_XS quantization for 16GB gang
- Show HN: Laptop is the last place your secrets are still in plaintext
- The Architect: Interactive Visualization of Deep Learning Mathematics Directly in Microsoft Excel
- Agentao: A Governed Local-First Runtime for Tool-Using LLM Agents
- Petition to add a rule for people to add their DAMN quant levels to their posts
- Ling 3.0 support merged into llama.cpp
- 100$ worth of gpu runs qwen 3.8 27b at 7.39 t/s
- Controlling Android with Gemini 3.7 Flash and 150 lines of Python
- After pushing 1M+ tokens through Qwen 3.8 27B, here is my optimal llama.cpp config for 16GB VRAM (73k Context, Agentic Coding)
- llama.cpp version v0.1.0 has been released
- Talk to Puck
- Same Cluster, 33 Points More Utilization: What Changed Was the Order
- llama.cpp adaptive MTP PR#27210
- Education Discount
- Qwen 3.8 27b saved me $650+ in API costs this evening
- Agentic Commerce at Scale: Your LangChain agents can transact securely
- Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers
- Running DeepSeek V4 Flash Q4_K_XL at ~100 tok/s prompt processing on 4× RTX 3060 12GB
- Introducing LangSmith Tuned Evaluators, starting with Perceived Error
- Asana cleared 5 years of engineering work in 2 weeks with Codex
- Claude Design artboard workflow added to Claude Code CLI
- Show HN: Runbook.v1 – governed workflow execution for MCP (fail-closed)
- Enterprise managed settings in GitHub Copilot for JetBrains
- Qwen3.8-27B on 2x 3090 + vLLM + DFlash2: 218 tok/s single request
- MCP in Orbs
- Replit expands access to software creation with GPT-5.6 Luna
- NVFP4 on VOLTA! Despite being built for Blackwell, I made four 2017 V100s run Qwen 3.8 NVFP4 natively and match my $6000 RTX 5090.
- Pass the Orb to the Left Hand Side
- DFlash2 speeds Qwen 3.8 27B up to 4 times
- Tool: TurboVec
- Beta: Saggar
- Beta: Needle
- DeepSeek Just Built the Next Generation of Coding Agents
- The boring way to run Deepseek V4 Flash-0731 130-150 tks - 16x5060ti 16GB over 2 PLX88096 switches
- TinySearch v0.6.1 - still a lightweight web research tool for local LLMs, now with bring-your-own-browser support
- How ChatGPT Work helps Stampli move ideas to market
- LangSmith Preview Builds: Test agent changes before production
- Unsloth Dynamic 3.0 GGUFs
- Show HN: Huzzah – a novel approach to coding with AI
- Qwen 3.8 27b - PI AGENT vs OPENCODE
- An Evidence-Grounded Multi-Agent System for High-Level Bio-Robot Design
- The Evaluation Context Protocol (ECP): A Portable Contract for AI Agent Evaluation
- Qwen3.8-27B at 262K context on a Strix Halo + RTX 3090 Ti: 9.5 -> 153 tok/s, and it beats a dual-3090 vLLM box on HumanEval
- Fastest NVFP4 quant of Qwen3.8 27B out there
- Explain Usage
- ChatGPT Apple Messages 💬, Anthropic’s meeting recorder 💼, Mistral Agentic Search 🔍
- DeepSeek Harness v0.1.1 released
- Filenames are the wrong index for Claude Code @ mentions
- Shared agentic work with GitHub Copilot in Microsoft Teams
- Getting the Same Results with Smaller "Cheaper" Dual Sparks AI as the More Expensive Clusters
- Qwen3.8-27B Q6 is a beast at agentic coding
- 16 GB VRAM purgatory discussion thread
- Show HN: OzBrain, a shared brain for knowledge between agents and your team
- Show HN: Omacosy – Omarchy-style tiling desktop for macOS, no SIP
- Show HN: Shoehorn – Quantize any model down to run on your machine
- This is why I run locally.
- Qwen 3.8 27b - PI AGENT vs OPENCODE - another smaple
- Llama.cpp version 0.2.0 is out!
- The New MCP Roadmap
- The Official Ruby SDK for MCP Reaches 1.0
- Show HN: Git workflow as an AI-agent skill
- I forked Ninfer 3090 and converted it to run on the CMP170HX - doubled my Qwen3.6-35B from llama.cpp
- GLM and I created a llama.cpp fork optimized for AMD GFX906 (Mi50, Mi60, Radeon VII, GCN HIP) - Machine Learning, LLMs, & AI
- I benchmark DFlash 2 (PR build) in llama.cpp on Qwen 3.8 27B against all speculative methods for 3 days. 2.26x on 100 real coding prompts, 4.68x with one n-gram drafter on top. Up to 8x on specific cases.
- Show HN: terminal-code – VS Code inside the terminal
- I fine tuned Gemma 4 12B for a 2.7x improvement on tool calling because I can't fit anything else comfortably into my 16 GBs of Vram
- I hosted Kimi K3 (2.8T parameters) using 8 B300s. 92 tok/s, $190 per million tokens
- i finally switched from windows to linux and got a 30-50% boost in speed.
- Show HN: 26-node n8n workflow for scoring and routing B2B leads
- Qwen3.5-9B Triple-Loop
- Qwen 3.8 27B for actual local programming
- Show HN: Froging AI – image and video models in one workflow
- Friendly URLs for Sharing Orbs
- Qwen 3.8 27b helped me with something unique that Opus 4 couldn't - Firmware + Software preservation and emulation on an early 2000's ARM based POS system
- Live Artifacts: Authoring Dynamic Media via Live Layers Encapsulating Generative Specifications
- PrimeAgentOrchestrator: Memory-Primed Agent Spawning for Personal AI Infrastructure
- I developed my own quantized LLM from scratch, trained on 30B tokens, deploys in 60 MB
- FreeToken Deepseek V4 Flash on a Single 3090 Local AI Testing
- Advancing price-performance for developers with GPT‑5.6 in Kiro
- JetBrains local AI (using Qwen3.6 27B)
- Please join r/LowEndLocalAI, a community for running local LLMs on low spec hardware
- Wire It, Run It, Deploy It: AI Workflows in Gradio
- I just tried DeepSeek Harness and it escaped from its workspace folder
- OptiMAS: Automatically Optimize Multi-Agent System
- Minimal Local Simulation Foundations for LLM- and VLM-Driven Agents in 2D and 3D Environments
- Show HN: Kern – container and resource runtime in a 1.5 MB binary, no daemon
- Do not blindly delete your older models, some are still precious
- Show HN: Screen memory without screenshots, just text to Markdown
- Every visual workflow tool becomes spaghetti. Here's what I built instead
- How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code
- New: Llama.cpp adaptive speculation for faster inference
- New in LangSmith Engine: >2x better issue detection
- Introducing the Admin plugin for ChatGPT Work and Codex
- Show HN: I made a Raspberry with Qwen my local car AI
- GitHub Copilot app Customize tab is generally available
- Maiao: Gerrit-style code review workflow for GitHub, GitLab, Gitea, others
- Setup Without a Commit
- Whoever the fuck predicted we would have gpt 5.5 performance in coding on consumer hardware a couple months ago now, i applaud you
- The Future of SaaS Is Apps That Agents Can Use
- Show HN: Rudder – Red-Green TDD Workflow for Verifiably Comprehensive Specs
- August 2026: LangChain Newsletter
- Lemonade end-of-summer project update, now serving 15 engines!
- Beta: Huzzah
- Beta: Kern
- Beta: Apache Maka
- Beta: OpenViking
- Enterprise-managed settings now support autoUpdate for plugin marketplaces
- MCP-Driven Accessibility Tree Standardization for AI-Powered Screen Reader Agents
- So Long, TUI Sidebar
- llama : add --n-cpu-ffn option by John-194 · Pull Request #26622 · ggml-org/llama.cpp
- Projects with Multiple Repositories
- We’re the Team Behind Apodex 1.1 — Ask Us Anything!
- Request: unsloth Please re-quantize Qwen3.6 35 A3B and 27B using UD 3.0
- No, Engrams won't let you run 1T models locally. It does something even better.
- Show HN: My Claude quota ran out in 10 minutes, so I made a tool to find out why
- llama.cpp support for Qwen3.8-Flash-Next has been merged
- Copilot code review: Resolution reasons and expanded capabilities
- Actions retention will cover checks, workflow runs, and statuses
- Show HN: We built open OpenRouter that turns usage into a better model
- Previewing the Model Hardware Standard
- Amp on iOS & macOS
- Actions retention will cover checks, workflow runs, and statuses
- It's unbelievable! I used the mmap function in llama.cpp to fit Qwen3.8-Flash-Next IQ3_XSS into 16G+64G RAM, and the speed still reached 26t/s.
- ds4 branch with GLM 5.3 Flash support
- Qwen3.8-Flash on RTX3090 + 64GB RAM (but you only need 12GB VRAM)
- No Mailmap Required
- Rime Finally Made Voice Agents Good Enough
- GitHub Copilot in Visual Studio — August update
- GitHub Copilot weekly releases — August 24
- A smarter way to run code migrations with less LLM context
- [Release] SOTA GGUFs for Qwen3.8-27B: GSQ-RCO at 2.5 to 3.0 bpw
- Is it worth running Qwen 3.8 Flash Next on 4x3090 vs 27B?
- Qwen 3.8 Flash Next ngram look up table offloaded to SSD and streamed in SGLang
- Qwen 3.8 27B at 50 tok/s with 100k Context on a 16GB GPU! (beellama.cpp)
- llama.cpp Open PRs list - CPU/RAM/Disk/Hybrid Related - Better for CPU-only & Hybrid inference
- Koboldcpp v1.120 released
- I fine-tuned a 0.8B local model for dictation cleanup. It matched a hosted frontier model on this narrow task
- Uncensored Multi-Model Releases, LongCat-Flash-Lite-Sparse with MTPs and LSAs, Qwen3.8-27B with MTPs, Qwen3.5-122B-A10B with MTPs, Qwen3-Coder-Next and Laguna-S2.1 with Vision, All Available in GGUF Format! Bonus: Links to my llama.cpp Fork for LongCat-Flash-Lite Support and J-Wash Enhanced Fork!
- Don't Sleep on EXL3 Quants
- Qwen3.8-Flash-Next turns 4xR9700 into a local AI powerhouse! 120 t/s TG and 12k t/s PP single request with optimized vLLM
- Demo of local document extraction (52 pages) using Arctic Embed and Bonsai 8B on an Iphone 16 (KernelAI app)
- Qwen 3.8 Flash Next locally on simple mobile phone at 3.5 tok/s
- Here my pretty good qwen3.8 27B setup, hope it helps
- GOD: Govern, Observe, and Direct - A Real-Time Control Room for Agent Societies
- pipecat-ai/phonellm-alpha-1: GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost
- How I got Qwen 3.8 27b running at ~75t/s decode on 16GB RTX 5080
- Setup OpenClaw 2.0 with Gemini in Under 60 Seconds
- First time running local models
- GitHub Copilot in VS Code, August 2026 releases
- AVX2: Speed up large batch size prompt processing of IQ models by bartowski1182 · Pull Request #27402 · ggml-org/llama.cpp
- Show HN: Silent Fail, get an email when your n8n workflow stops running
- The Web-CLI: Verifiable Privacy for Tools, Models, and Inference Engines in the Browser
- Mac ← USB-C cable → Linux box is becoming a thing.
- ExLlamav3 Recent Updates : CPU offload, GLM-5.3-FLASH, Qwen3.8-Flash, SC Quants ++
- I pushed Qwen3.8-27B to 2.000 prefill per second and 132 decode per second on A RTX 3090.
- Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI
- 11 Tiny Coding Agent Fixes With A Stupid Amount Of Payoff
- Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
- Intelligently Ordered Diffs
- Fable 5.1
- Selected GitHub Copilot models deprecated
- Harness Engineering: Anatomy, Architecture, and Evolution of Coding Agents -- A Source-Code Study of Eleven Systems
- CUDA-Harness: Harnessing Agentic CUDA Kernel Generation and Optimization from Natural Language
- ChatDev 2.0: A No-Code Multi-Agent Platform for Developing Everything
- Running 104GB Qwen3.8-Flash-Next on 48GB Mac at ~12 tok/s
Papers & ergonomics
- Central Tendency Bias in Human Selection of AI-Generated Design Variations
- Knowledge-Based Design Requirements for Generative Social Robots in Higher Education
- Advances in footwear and surface for prevention of occupational slips and falls - A scoping review
- Analysis of work system components in interprofessional communication to determine shock etiology
- The distracting role of stress: Impaired executive attention and delayed fatigue perception
- Vision-language models for occupational physical exposure assessment: Classification and temporal segmentation of manual material handling tasks
- Digital tools for identifying work tasks in various occupations: A scoping review
- Contributions of headset IPD fit, vection and sway to cybersickness during head mounted display based virtual reality
- Neurophysiological synchrony as an emergent performance marker of multi-human–robot team effectiveness
- A model for predicting pointing time in an eye-gaze input system using three basic phases of cursor movement trajectories
- Estimation of dynamic spinal loads during manual lifting using smartphone-based markerless motion capture
- Evaluating model-estimated shoulder muscle activity during overhead work with varied task demands and exoskeleton use
- Effects of sleep deprivation on cognitive-motor functions and adaptive skill learning among medical residents across 26h night shifts
- Population-specific ergonomic design and evaluation of a head-mounted display based on Chinese craniofacial measurements
- Import AI 456: RSI and economic growth; radical optionality for AI regulation; and a neural computer
- Effects of static postural loading on the performance of short-term/working memory tasks
- Gender and body height discriminate spinal movement patterns during lifting and lowering tasks
- Switching between touch and voice: factors influencing modality selection in multimodal systems
- Impact of anthropomorphism in AI assistants’ verbal feedback on task performance and emotional experience
- Evaluating biomechanical risks in manual material handling: an ergonomic intervention approach
- Effects of passive Arm-support exoskeleton on dynamic balance in different occupational tasks
- Learning to apply Design Thinking in participatory ergonomics: an exploratory study of OHS professionals in Denmark
- Recognising and explaining mental workload using low-interference method by fusing speech, ECG and eye tracking signals during simulated flight
- Size matters ̶ effect of screen setup on muscle activity and posture in computer work
- Leveraging socio-technical systems to tackle grand challenges: Reflections on human-robot teams, hybrid workplaces, med-tech, and digital transformation
- My coworker's 36 key Corne open-source keyboard setup
- When LLM Tutoring Responses Work: Evidence from Student Programming Conversations
- Same Stories, Different Journeys: From Social Comparison to Sensemaking in AI-Mediated Peer Career Exploration
- HandPad: A Bimanual Hand Interface for Fluid Window Interactions in VR
- Supporting Reflection in LLM-based Exploratory Search
- Faster AI, Uneven Frontier: Rapid Crossings, a Jagged Frontier, and the Repositioning of Human Judgment
- TRAIL: A Platform for Configurable Human--AI Teaming Experiments
- AI advice suppresses people's willingness to say "I don't know", even when the advice is wrong and accuracy is incentivized
- Role-specific experiences of inefficiency in the biomedical research and clinical laboratory workforce
- Conversational Tactile Data Interfaces: Co-Designing Accessible Data Experiences with Blind Users Using Refreshable Tactile Displays and Conversational AI
- When AI Blurs the Boundaries of Contribution: An Empirical Study of Authorship Calibration
- Measuring How Students Rely on Generative AI in Academic Writing: Development and Multi-Source Validation of the Generative AI Reliance Types Scale (GenAI-RTS)
- Can AR Embedded Visualizations Foster Appropriate Reliance on AI in Spatial Decision-Making? A Comparative Study of AR X-Ray vs. 2D Minimap
- Initial static, pseudo-static, dynamic, and cognitive fit evaluation of three passive shoulder exoskeletons while performing simulated manufacturing tasks
- Human factors integration in complex systems: Awareness, challenges and strategies
- Smart glasses: environmental perception and perceived health effects
- System for activity-aware fatigue evaluation (SAFE) framework: Predictive fatigue modelling for occupational tasks
- It Matters How You Say It: Exploring Rhetorical Patterns for AI-Assisted Information Evaluation
- Informal Learning Emerges in Everyday Human-LLM Interaction
- Effectiveness of passive back-support exoskeletons during simulated commercial crab fishing tasks: Acute effects on muscle activity, joint kinematics, and subjective measures
- Effects of uncertain system information on users' performance, confidence, and trust in aided linguistic judgement
- Toward context-aware and personalized sit-stand desk interventions: Insights from a field observational study
- Textual cues, cognitive load, and social fatigue: Unveiling the reasons behind user discontinuance in conversational AI
- Knowledge Gaps and Explanation Design in AI Advice: A Cognitive Fit Perspective on User Decision Confidence
- The creative tax of videoconferencing: How ideological diversity mitigates creativity losses in virtual teams
- Trust erodes, fatigue builds: How prompt uncertainty traps users in recrafting loops
- The Digital Mindfulness Scale: Development and longitudinal validation in the workplace
- A Comprehensive Evaluation of Job Rotation: Biomechanical Risk, Body Discomfort, and Psychosocial Demands
- Comparing Paper- and AR-Based Assembly Manuals in Task Performance, dlPFC Hemodynamic Responses, and Perceived Workload
- The Influence of Exposure and Error Type on Estimates of Automation Reliability
- Mind-Wandering or Task-Unrelated Thought Reports May Be a Response to Performance Not a Cause of Performance: Using Forced Errors to Impact Thought Content Reports
- Vigilance Research Beyond the Laboratory: Methodological Considerations and Practical Insights
- Efficiency Pitfalls of Explainable AI in Clinical Diagnostic and Treatment Human-AI Workflows
- Human–Artificial Intelligence Collaborative Decision-Making in Emergencies: Relative Advantage Theory
- Modeling Human Expertise in a Sanding Task
- Compatibility Effects With Simple Lever Tools: A Replication and Extension Beyond Simple Button Responses
- System-Wide Trust (SWT) Versus Component-Specific Trust (CST) in Multi-Agent Human–Agent Teams: Individual Variability in Trust Bias
- Effects of Task Priority and Difficulty in Multitasking Across Screens
- Ethical Decision-Making in the Workplace: Which Factors Influence Decision-Makers’ Willingness to Revise Their Decisions due to AI-Generated Suggestions?
- Decision-Supported Visual Search: Effects of Direct and Indirect Cues on Performance and Attentional Mechanisms
- A Structural Model of Attentional Effort Dynamics: Evidence From a Naturalistic Discrimination Task
- Affective Interruptions Impair Task Performance Less Than Neutral Ones Across Age and Task Demands
- Comprehensive Evaluation of Explanation Types in a Spaceflight-Relevant Human–Autonomy Teaming Task
- What’s in a Name? Implications of AI Roles and Mind Perception for Human-AI Teams
- Can We Learn From Trust in Simulation to Gain Trust in AI?
- Systemic Trust in Artificial Intelligence
- Does Expertise Matter? A Study of Chile’s Ergonomic Risk Assessment Tool
- Effectiveness of Eye-Tracking Metrics in Human-Centric Design of Human-Machine Interface: Cases on Process Control Operations
- Three Participatory Methods to Engage Employees in Workplace Research and Design
- An Exploratory Study of the Causes of Discomfort From Using Braille Displays
- Spatial Discontiguity in Three-Dimensional Augmented Reality Spaces
- Satellite Ground Stations: The Need for Additional Research and Standardization
- An Empirical Study of Tablet Ergonomics: The Interplay of Temperature, Orientation, and Use Behaviors
- Contours of Comfort: Mapping Pressure Landscapes Across Anthropometric Seating Interfaces
- Workplace Ergonomics in Bangladesh: A Scoping Review of Anthropometric Mismatches and Musculoskeletal Disorders
- Querying Multimodal Scientific Papers with AI: Practices and Preferences Across Blind, Low-Vision, and Sighted Scientists
- TargetFinder: Detecting Widgets from Pixels on Desktop Interfaces
- Magnus Evo XL review: Height-adjustable gaming desk put to the test
- Mitigating psychological discomfort during automation error: A mixed methods research
- CRAFT: Exploring Wearable Creative AI on Smart Glasses for Fiction Writing in Real-World Contexts
- Transparent by Design, Usable in Practice? A Formative Usability Study of a Conversational Product Advisor
- Delegating (or Not) to Machines: How Role Expectations Shape Leaders’ Willingness to Use AI for Communication
- Bespoke Visual Assistance: What and How do Blind and Low-Vision People Create with Agentic Programming?
- A Method for Constructing a Dynamic Seat Adjustment Model to Mitigate Prolonged-Sitting Discomfort
- Touching or Chatting: The Utility of LLMs and Tactile Charts for Learning about Complex Chart Types by BLV Individuals
- Relationships Between Trust, Compliance, and Performance for Novice Programmers Using AI Code Generation
- Verification Without Distrust: Reframing User-Side Oversight as Routine Epistemic Governance in Everyday Human-Chatbot Interaction
- Using Eye Tracking to Identify Comfort Attributes in Office Chairs
- How Wrangling Tools Shape Wrangling: A Technical Dimensions Analysis
- FleetScape: A Mixed Reality Sandtable for Spatial Supervision and Control of Scalable Drone Fleets
- Evaluating the Vergence-Accommodation Conflict in Gaze-Based 3D Target Selection
- Stop Shipping AI Agents on Faith: Capability Is Not Production Readiness
- $\Sigma$-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems
- More externalization, but less inference? Exploring changes in young learners’ critical thinking during conversational AI bot-supported multimodal writing practice
- An experimental investigation on the relationship between seat pan and seat back angles to eliminate the shear force on the seat pan
- Cognitive Readiness for Human-AI Collaboration
- Visualizing Placement Proposals for Window Arrangement in Mixed Reality: A Comparative User Study
- Chat Debugging: An Exploratory Study of Human-AI Collaboration to Debug Analog Circuits
- Efficient Optimal Mouse Sensor Position Estimation using Simulated Cursor Trajectories
- Revisiting Channel Effectiveness: A Multi-Dimensional Evaluation with Primitive Visual Stimuli
- Mixed Uncertainty in One View: Co-Visualizing Statistical Variability and Qualitative Confidence
- Toward Resilient Human-AI Collaboration: A Lifecycle Taxonomy of Sociotechnical Risks and Cascading Failures
- CaRing: Preventing Carpal Tunnel Syndrome based on Daily Activities from Always-Available Input Device
- Topic Matters: How Linguistic Properties can Shape Reading Behaviour in Selective Exposure Studies
- Obesity-related differences in shoulder and lumbar loading during manual material handling
- Effect of non-neutral seated postures on whole-body vibration measurements
- Not Always Top-Left: Untangling the Signals that Guide Dashboard Reading Order
- Ergonomic office chair with lumbar support and footrest: Welax S9 Pro hands-on review
- Large Language Models Explain Experts Better Than Experts Themselves
- Concurrent validity of markerless motion capture for occupational lifting: Kinematics and ergonomic risk assessment
- Navigation Alone Is Not Enough: Evaluating Explanatory Assistive UI Agents
- TRIBE: Predicting Team Performance via Communication Behavior Ensembles
- How Do Senior Qualitative Researchers Perceive the Risks and Opportunities of Using Large Language Models in Qualitative Analysis? An Exploratory Study
- Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence
- From “cost of asking” to “fit of asking”: How seeking help with AI shapes employees’ indebtedness and autonomy in different workplace helping contexts
- Ontology-Grounded World Models for Failure Diagnosis and Closed-Loop Repair in Physical AI Systems
- How Cognitive Load Affects Dynamic Trust Calibration in Human–AI Collaboration: Evidence for Selective Pathway Effects
- YouthShape.US: A New Anthropometric Resource for U.S. Children and Youth
- What Cognitive Accessibility Reveals About Data Visualization
- Spatial alignment of audio–haptic directional feedback shapes non-visual hand guidance
- Chat First, Worry Later: Understanding Individuals' Privacy Perceptions Using ChatGPT in a Work Context
- What is the current state of evidence on passive exoskeletons designed to support the lumbar region? A systematic review and meta-analysis of physiological, biomechanical, and user-related outcomes
- Assessment of pressure discomfort for over-ear wearables and the relationships with objective metrics
- ‘I’d feel like management understands (no pun intended) how we feel’: evaluating a hypothetical policy promoting sitting in standing-biased jobs
- Effects of prismatic loupes on surgeons’ intraoperative physical workload and musculoskeletal discomfort in operating room
- Validating force-estimating insoles for calculating centre of pressure and vertical ground reaction forces during occupational tasks
- Mastering a robot workforce: review of single human multiple robots systems and their impact on occupational safety and health and system performance
- Automatic estimation of Hand Activity Level from upper-limb trajectories: a probabilistic regression framework
- State of science: New frontiers in inclusive design and digital health intervention
- Physiological measurement of situation awareness: a study of the validity of EEG and fNIRS during performance and automation monitoring in a complex task
- Effects of dynamic lighting on neurobehavioral performance under different mental states in the working area of a space station
- Aura: Dynamic Intra-Turn Emotion-Aware Adaptation of Large Language Model Responses
- Using physiological measures to assess and diagnose team performance: An interrogative approach
- Investigating Human Factors and Ergonomics research: a 4S framework
- Applying cognitive and perceptual science to typeface choices
- A scoping review on emerging technologies and automation of musculoskeletal ergonomic assessments
- Strategic music listening and subjective perceptions: effects on attention and workload across task complexity levels
- Enhancing mental workload recognition: a comparison of complexity-based eye movement metrics and conventional features
- An ergonomic intervention to minimise physical and physiological stresses in the office standing workstation
- REBA integrated with organisational analysis to assess the risk of biomechanical overload in physiotherapists
- The effects of prolonged standing in occupational footwear on perceived discomfort, standing balance, and gait biomechanics in young adults
- RegulAR: Graph-Grounded Error Recognition and Assistance for Procedural Tasks in AR
- Dynamic Tree Colors: Adaptive Discriminable Hierarchies with Minimum Instability
- Graphionale: How Graph Visualizations of LLM Rationales Affect Human Decision Making
- Too Much of the Same: From Algorithmic to Human Bias in Learning to Defer
- User Preferences for UI Anchoring in MR: Effects of Task Mobility and Interface Properties
- Not All Explanations Are Sought: Information-Seeking Psychology for Human-Centered XAI
- Toward Postural State Classification in Immersive VR with Multimodal Data and Explainability Analysis
- Collaboratively Eliciting Gestures for Geospatial Data Exploration on an MSE with Tangibles and Styluses
- FocusBuddy: Encouraging Healthy Desk-Work Habits by Caring for a Virtual Pet on a Water Bottle
- ErgoAssist: Cognition-Aware Posture Feedback in Wearable Ergonomic Systems
- RecalibrateGPT: AI Fatigue Resilient Conversational Interfaces
Practical tips
- Giving your AI a Job Interview
- Layout Buffet - Mousing with a keyboard
- Layout Buffet - Layers
- Practice typing with your favorite book excerpts
- Layout Buffet - Home-row Mods
- Layout Buffet - Sticker Mods
- Automate Excel with Python: From manual grind to one-click workflow
- Brave's latest browser release offers Containers for better and easier workflow
- Made a free macOS menu bar app that fixes typing in the wrong keyboard layout
- Mouseless – keyboard-driven control of macOS/Linux/Windows
- Neverclick: Desktop application for performing mouse actions with your keyboard
- Getting started with ChatGPT
- Gemini 3 Prompting: Best Practices for General Usage
- Hidden Gems: Part 3
- Show HN: I was tired of opening 2 tabs for every HN link, so I made a userscript
- PSA for DeepSeek-V4-Flash-0731 users — don't blow out your prompt cache with system role messages mid-conversation
- You really should not quantize KV Cache for DeepSeek V4 Flash
- How to Decide When an AI Tool Is Worth Keeping
- My Terminal Workflow for Note-Taking, Data Engineering and Writing (Linux/macOS)
- One AI Output Is an Example, Not an Evaluation
- Weirdly, no one talks about Temperature setting for the Qwen3.8 27b
- GPU Poor - Don't overlook Laguna XS 2.1
- Don't sleep on Vision support for coding!
- Everyone is t/s maxing.. 3.8.. but after a week of using it for work I'm tempted to switch back to 3.6
Raw data: signal-graph.json
Signal. An evidence stream on how AI is changing the workday — sources read, scored, and filed.
read the stream →- Score 3.80curatedAgent tools & setup
ChatDev 2.0: A No-Code Multi-Agent Platform for Developing Everything
ChatDev 2.0 (DevAll) is a no-code platform for building, executing, and inspecting LLM-based multi-agent systems. It pairs a declarative executable graph with a cycle-aware execution engine to support heterogeneous agents and cyclic interactions, and provides a visual interface for authoring and monitoring without code. Open-sourced on GitHub by OpenBMB.
arXiv cs.MA · T1
- Score 3.80curatedAgentic workflow patterns
Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs
The paper proposes control-data flow separation for multi-agent LLM systems: execution-critical protocols become typed, validated program objects while task content remains optimizable natural language, preventing prompt edits from corrupting routing or formatting logic. Tested on reasoning, review, and insurance rating workflows with 100% protocol validity.
arXiv cs.MA · T1
- Score 4.40curatedAgentic workflow patterns
Long-Horizon State Tracking in LLMs: Executing MD5 through a Deep Sequence of Dependent Tool Calls
A research paper tests whether LLMs can maintain exact intermediate state across long sequences of dependent tool calls by having a model compute MD5 step by step across 196 calls and 64 rounds. Using gpt-oss-120b, it finds that keeping the model's own reasoning in context and voting over a thinking-enabled worker enables correct end-to-end execution.
arXiv cs.MA · T1
The ledger’s credibility
How the Signal ledger earns trust.
Four rules keep every reading auditable — from where a claim comes from to how the map is allowed to draw it.
Tiered, reviewed sources
Every source sits on a reviewed list. Its tier weights the final score by provenance — an auditable code weight, not a trust badge.
Computed, not conjured
Five dimensions are scored, then combined by formula — mean × source weight. The final always says what it is, never a model’s opinion.
Source and reading kept apart
What a source said and what the pipeline read never merge into one voice. Pipeline notes always carry their disclaimer.
Machine guesses stay labelled
On the map a dashed link is an embedding’s suggestion; a solid link is an editor’s citation. The two never blur.
Tile values are illustrative examples — not a live entry’s reading.











