explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR: the questions people are asking
  • The problem: reaching files is not the same as using them
  • What they measured
  • How RunningTab works
  • Why environment-side beats model-side notes
  • How this differs from agent memory and planning
  • How to copy the pattern in your own agent
  • Limits and open questions
  • Why this matters
  • Related reading
← Back to blog

explainx / blog

RunningTab: Why File-Reading Agents Drop Facts They Already Saw, and a Fix

AI Research, AI Agents, Agent Memory, Context Window, Knowledge Work

RunningTab keeps an environment-side record of what a task owes and what an agent has read. How it works, the failure it targets, and what to copy today.

Oct 8, 2026·8 min read·Yash Thakker
add explainx.ai
go deep
RunningTab: Why File-Reading Agents Drop Facts They Already Saw, and a Fix

Agents that work through a folder of files have a quiet failure: they open the right document, see the right number, and still leave it out of the final report. A paper trending on Hugging Face this week, RunningTab (arXiv 2610.10444), measures that failure and proposes a fix that does not touch the model at all. The environment keeps a running tab of what the task still owes and what the agent has already read.

This post explains the idea in plain terms, what the authors measured, how it compares with the usual memory and planning tricks, and how you could copy the pattern in your own agent harness today.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

A green lens examining a cream box with a short branching path, illustrating auditing of what an agent read and delivered

TL;DR: the questions people are asking

table · 2 cols
QuestionShort answer
What is it?A framework that records, outside the model, task requirements, files read, and files listed but never opened
Who wrote it?Jinheon Baek, Soyeong Jeong, Yumin Choi, Dongsu Han, Sung Ju Hwang; submitted to arXiv October 7, 2026
What failure does it target?Content the agent saw but did not deliver
Does it need a new model?No; it changes the environment, not the weights
How was it tested?Three benchmarks, three LLMs; it beat plain direct workspace interaction and model-kept-notes baselines in every case
Is code available?Not stated in the sources we reviewed

The problem: reaching files is not the same as using them

Knowledge work often means producing a new deliverable, such as a report, a summary, or a spreadsheet, from files a workspace already holds. Agents increasingly do this with a terminal: they list directories, grep, open files, and write the result. The paper calls producing a deliverable this way direct workspace interaction (DWI), building on earlier direct corpus interaction work that skips indexing and lets the agent search raw files.

Most prior diagnosis of failures in these benchmarks blames reaching or understanding files, for example not understanding documents or not tracing lineage across files. The RunningTab authors argue there is a second loss that those diagnoses miss: content from files the agent had already opened, which showed up in its observations but never made it into the output.

What they measured

To size the problem, the authors analyzed plain DWI (called DCI) running GPT-5.4 nano on Workspace-Bench, a benchmark that scores deliverables against rubrics, across three runs of its 100 tasks. They labeled each failed rubric with LLM annotators, who agreed with human labels at a Cohen's kappa of 0.83.

The breakdown of lost points is striking:

  • Extraction failures were 51.1% of lost points. The agent opened the right files but the deliverable lacked or distorted the content.
  • Within that, Shown but Missing was 12.3% (the content appeared in an opened file's output yet was absent from the deliverable) and Dropped or Distorted was 38.9%.
  • Discovery failures were 17.6%, where the right files were never opened, including files that were listed but not opened.
  • Among attempts that produced a deliverable, 19.8% of the 891 distinctive numeric values that rubrics expected and that appeared in an observation were missing from the deliverable.

In other words, about one in five relevant numbers the agent had literally seen did not make it into the final answer. That is a context-window problem. Observations pile up across dozens of turns and old ones stop influencing the output. For background on why, see our context window explainer and Claude Code context management guide.

How RunningTab works

The core idea is a per-task tab, kept by the environment next to the agent. It has three kinds of entries:

  1. Requirements. The agent states what the task asks for, as a list of items the deliverable must contain.
  2. A read log. Every file the agent reads is recorded by the environment as an excerpt with its provenance, meaning where in which file it came from. The agent does not have to remember to save anything.
  3. Candidates. Every file that appeared in a listing but was never opened is tracked as an unopened candidate.

During the run, the agent can request a review that shows each open requirement beside its best-matching excerpts and the top unopened candidates. It then resolves a requirement by citing the read-log entry that holds the answer, or sets it aside with a stated reason. If the agent tries to finish while requirements are still open, the environment returns a single finish check listing what is unresolved.

In the paper's worked example (Figure 3), the missing fact had already been captured by the tab, so at the finish check the agent simply closes the requirement against the logged excerpt and delivers.

Why environment-side beats model-side notes

A natural alternative is to ask the model to take notes or maintain a plan. The paper compares against baselines where the model itself keeps the record. Two things make the environment version more reliable:

  • Capture is automatic. The environment logs every read, so nothing depends on the model deciding a passage is worth saving. Ablations in the paper attribute most of the gain to what the environment captures on its own.
  • The finish check is a hard stop. A model-written plan can be ignored. A check that returns open items at the moment of submission is hard to skip.

The authors also report that the tab usually holds the values a deliverable needs once the agent has seen them, and that the gain grows with how much the agent reads. That last point matters for practice: the more files in play, the more a context window leaks, and the more a tab helps.

How this differs from agent memory and planning

Agent memory research mostly targets what survives across sessions: reflections, stored experiences, retrieved notes. Planning work targets what the model writes into its own context within a task. RunningTab sits between them. It is within-task state, but owned by the environment and indexed against the task's requirements and the files' provenance. It also changes nothing about how the agent reaches files. That is why it can bolt onto an existing agent loop, and why it is complementary to retrieval ideas such as agentic search as code and to memory patterns like the LLM wiki approach.

How to copy the pattern in your own agent

You do not need the paper's code to apply the idea. A minimal version looks like this:

  1. Add a requirements tool. Let the agent register items the output must contain, and require it to do this first.
  2. Wrap your file-reading tools. Every read appends an entry (file path, line or page anchor, excerpt) to a log the harness owns.
  3. Track listings. When a directory listing or search returns files, mark them as unopened candidates until read.
  4. Add a review tool. Show each open requirement with the best-matching log entries, using simple keyword or embedding match.
  5. Gate submission. If requirements are unresolved, return them instead of accepting the final answer, and require a reason to set one aside.

This costs a few tool definitions and some storage, and it turns "did the report include the Q3 figure we saw?" from a hope into a checkable fact. It fits the broader harness engineering lesson that environment design often moves results more than prompts do, and it is directly relevant to anyone building agents for knowledge work.

Limits and open questions

  • We have only partial numbers. The abstract and introduction say RunningTab beat plain DWI and model-kept baselines in every case across three benchmarks and three LLMs, but we did not extract per-benchmark scores. Check the paper's tables before quoting a gain.
  • Measured failure analysis used a small model. The 19.8% figure comes from GPT-5.4 nano on Workspace-Bench. Stronger models may lose less, though the authors' framing suggests the issue grows with the amount read.
  • It is a preprint. The paper was submitted October 7, 2026 and has not been through peer review, as far as we can tell. Hugging Face lists it among the day's trending papers.
  • Matching quality matters. The review step pairs requirements with excerpts, so weak matching could surface the wrong evidence. The finish check helps but does not guarantee correctness of the content itself.
  • Not a verification tool. The tab ensures content is not forgotten, not that the source is right.

Why this matters

Much of the current conversation about agent reliability focuses on bigger models and longer contexts. RunningTab is a reminder that a large share of failures is bookkeeping: the agent had the information and lost it. Bookkeeping is a solvable engineering problem, and solving it in the environment makes the fix portable across models. If you run agents over document folders, contracts, research corpora, or data rooms, a requirements-and-excerpts ledger is one of the cheaper reliability upgrades available.

Related reading

  • What is an agent harness? A complete guide
  • Agent harness engineering and Terminal-Bench
  • LLM context window explained
  • Claude Code context window limit management
  • Karpathy's LLM wiki pattern for agent memory
  • Perplexity search as code: agentic search
  • AI agents for knowledge workers

Primary source: the arXiv abstract page and the Hugging Face paper page.

Paper details are accurate as of October 8, 2026 and may change in later versions.

Spotted something out of date? Let us know.

People in this article

  • Andrej Karpathy →AI researcher and educator
Explore people in AI →
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Oct 7, 2026

Et Tu, Brute? Study Finds AI Agents Steer Wealthier Users to Pricier Flights, Insurance and Schools

Researchers from Cisco Foundation AI and Carnegie Mellon report that personal AI agents infer a user's wealth from email and profile data, then recommend more expensive flights, insurance and PhD programs, without being told to. Here is what the paper measured, what the caveats are, and what to do before giving an agent your inbox.

Sep 28, 2026

Thinking Fast and Slow in AI (2021): Metacognition on Hacker News Again

arXiv:2110.01834 (October 2021) describes a multi-agent architecture where fast System 1 solvers react from experience and slow System 2 agents deliberate when needed — backed by world and self models. A September 28, 2026 Hacker News thread asked whether GPT adaptive reasoning already "solved" this or whether Kahneman's psychology is the wrong blueprint. explainx.ai unpacks the paper, the replication debate, and what builders should steal for 2026 agent stacks.

Sep 10, 2026

The Second Writer: How Self-Evolving Coding Agents Actually Learn

Every coding agent is a frozen model plus a program that decides what it sees this turn. "Self-evolving" just means something else is now allowed to edit that program's inputs after the turn ends. mem0 mapped three very different versions of that edit onto one confusing number — here's the breakdown, and how to implement the persistent one with real memory scoping.