Agents that work through a folder of files have a quiet failure: they open the right document, see the right number, and still leave it out of the final report. A paper trending on Hugging Face this week, RunningTab (arXiv 2610.10444), measures that failure and proposes a fix that does not touch the model at all. The environment keeps a running tab of what the task still owes and what the agent has already read.
This post explains the idea in plain terms, what the authors measured, how it compares with the usual memory and planning tricks, and how you could copy the pattern in your own agent harness today.

TL;DR: the questions people are asking
| Question | Short answer |
|---|---|
| What is it? | A framework that records, outside the model, task requirements, files read, and files listed but never opened |
| Who wrote it? | Jinheon Baek, Soyeong Jeong, Yumin Choi, Dongsu Han, Sung Ju Hwang; submitted to arXiv October 7, 2026 |
| What failure does it target? | Content the agent saw but did not deliver |
| Does it need a new model? | No; it changes the environment, not the weights |
| How was it tested? | Three benchmarks, three LLMs; it beat plain direct workspace interaction and model-kept-notes baselines in every case |
| Is code available? | Not stated in the sources we reviewed |
The problem: reaching files is not the same as using them
Knowledge work often means producing a new deliverable, such as a report, a summary, or a spreadsheet, from files a workspace already holds. Agents increasingly do this with a terminal: they list directories, grep, open files, and write the result. The paper calls producing a deliverable this way direct workspace interaction (DWI), building on earlier direct corpus interaction work that skips indexing and lets the agent search raw files.
Most prior diagnosis of failures in these benchmarks blames reaching or understanding files, for example not understanding documents or not tracing lineage across files. The RunningTab authors argue there is a second loss that those diagnoses miss: content from files the agent had already opened, which showed up in its observations but never made it into the output.
What they measured
To size the problem, the authors analyzed plain DWI (called DCI) running GPT-5.4 nano on Workspace-Bench, a benchmark that scores deliverables against rubrics, across three runs of its 100 tasks. They labeled each failed rubric with LLM annotators, who agreed with human labels at a Cohen's kappa of 0.83.
The breakdown of lost points is striking:
- Extraction failures were 51.1% of lost points. The agent opened the right files but the deliverable lacked or distorted the content.
- Within that, Shown but Missing was 12.3% (the content appeared in an opened file's output yet was absent from the deliverable) and Dropped or Distorted was 38.9%.
- Discovery failures were 17.6%, where the right files were never opened, including files that were listed but not opened.
- Among attempts that produced a deliverable, 19.8% of the 891 distinctive numeric values that rubrics expected and that appeared in an observation were missing from the deliverable.
In other words, about one in five relevant numbers the agent had literally seen did not make it into the final answer. That is a context-window problem. Observations pile up across dozens of turns and old ones stop influencing the output. For background on why, see our context window explainer and Claude Code context management guide.
How RunningTab works
The core idea is a per-task tab, kept by the environment next to the agent. It has three kinds of entries:
- Requirements. The agent states what the task asks for, as a list of items the deliverable must contain.
- A read log. Every file the agent reads is recorded by the environment as an excerpt with its provenance, meaning where in which file it came from. The agent does not have to remember to save anything.
- Candidates. Every file that appeared in a listing but was never opened is tracked as an unopened candidate.
During the run, the agent can request a review that shows each open requirement beside its best-matching excerpts and the top unopened candidates. It then resolves a requirement by citing the read-log entry that holds the answer, or sets it aside with a stated reason. If the agent tries to finish while requirements are still open, the environment returns a single finish check listing what is unresolved.
In the paper's worked example (Figure 3), the missing fact had already been captured by the tab, so at the finish check the agent simply closes the requirement against the logged excerpt and delivers.
Why environment-side beats model-side notes
A natural alternative is to ask the model to take notes or maintain a plan. The paper compares against baselines where the model itself keeps the record. Two things make the environment version more reliable:
- Capture is automatic. The environment logs every read, so nothing depends on the model deciding a passage is worth saving. Ablations in the paper attribute most of the gain to what the environment captures on its own.
- The finish check is a hard stop. A model-written plan can be ignored. A check that returns open items at the moment of submission is hard to skip.
The authors also report that the tab usually holds the values a deliverable needs once the agent has seen them, and that the gain grows with how much the agent reads. That last point matters for practice: the more files in play, the more a context window leaks, and the more a tab helps.
How this differs from agent memory and planning
Agent memory research mostly targets what survives across sessions: reflections, stored experiences, retrieved notes. Planning work targets what the model writes into its own context within a task. RunningTab sits between them. It is within-task state, but owned by the environment and indexed against the task's requirements and the files' provenance. It also changes nothing about how the agent reaches files. That is why it can bolt onto an existing agent loop, and why it is complementary to retrieval ideas such as agentic search as code and to memory patterns like the LLM wiki approach.
How to copy the pattern in your own agent
You do not need the paper's code to apply the idea. A minimal version looks like this:
- Add a requirements tool. Let the agent register items the output must contain, and require it to do this first.
- Wrap your file-reading tools. Every read appends an entry (file path, line or page anchor, excerpt) to a log the harness owns.
- Track listings. When a directory listing or search returns files, mark them as unopened candidates until read.
- Add a review tool. Show each open requirement with the best-matching log entries, using simple keyword or embedding match.
- Gate submission. If requirements are unresolved, return them instead of accepting the final answer, and require a reason to set one aside.
This costs a few tool definitions and some storage, and it turns "did the report include the Q3 figure we saw?" from a hope into a checkable fact. It fits the broader harness engineering lesson that environment design often moves results more than prompts do, and it is directly relevant to anyone building agents for knowledge work.
Limits and open questions
- We have only partial numbers. The abstract and introduction say RunningTab beat plain DWI and model-kept baselines in every case across three benchmarks and three LLMs, but we did not extract per-benchmark scores. Check the paper's tables before quoting a gain.
- Measured failure analysis used a small model. The 19.8% figure comes from GPT-5.4 nano on Workspace-Bench. Stronger models may lose less, though the authors' framing suggests the issue grows with the amount read.
- It is a preprint. The paper was submitted October 7, 2026 and has not been through peer review, as far as we can tell. Hugging Face lists it among the day's trending papers.
- Matching quality matters. The review step pairs requirements with excerpts, so weak matching could surface the wrong evidence. The finish check helps but does not guarantee correctness of the content itself.
- Not a verification tool. The tab ensures content is not forgotten, not that the source is right.
Why this matters
Much of the current conversation about agent reliability focuses on bigger models and longer contexts. RunningTab is a reminder that a large share of failures is bookkeeping: the agent had the information and lost it. Bookkeeping is a solvable engineering problem, and solving it in the environment makes the fix portable across models. If you run agents over document folders, contracts, research corpora, or data rooms, a requirements-and-excerpts ledger is one of the cheaper reliability upgrades available.
Related reading
- What is an agent harness? A complete guide
- Agent harness engineering and Terminal-Bench
- LLM context window explained
- Claude Code context window limit management
- Karpathy's LLM wiki pattern for agent memory
- Perplexity search as code: agentic search
- AI agents for knowledge workers
Primary source: the arXiv abstract page and the Hugging Face paper page.
Paper details are accurate as of October 8, 2026 and may change in later versions.
