PERSONAL RESEARCH / Context & memory

Context engineering.
Working memory.

I explore how instructions, live state and persistent notes help agents continue longer tasks without treating the context window as a database.

Article date
Base notes dated
Sources checked
DENIZ WETZ / Personal research synthesisALL RESEARCH

01RESEARCH NOTE

Select the working context

Context engineering concerns the information available at each model call: instructions, tools, retrieved evidence and relevant history. Sourcegraph separates instructions, retrieval, memory and tools as a practical organizing model. The context therefore needs to be assembled for the current task. [01]

My working model separates stable instructions, authoritative live records, a small working-state note and knowledge retrieved on demand. I would read the current record through a tool, keep open questions and decisions in the working note, and load supporting documents when needed. This is my design synthesis rather than a required standard.

02RESEARCH NOTE

Persist the work outside the window

Anthropic describes compaction and structured note-taking for longer tasks. Compaction condenses a conversation; persistent notes can be read back into a later context. The tradeoff is what the summary retains: excessive compression can discard details that matter later. [02]

I want the saved state to include decisions, unresolved questions, the next action and links to actual outputs. A summary of an answer is not the saved result itself. I would keep source references with the result and reload the current state when work resumes.

Markdown and JSON are formats; PostgreSQL is a storage option; RAG is a way of retrieving evidence. They can work together. The choice depends on how information changes, how it is queried and what needs an authoritative record.

03RESEARCH NOTE

Diagnose the failure before changing the system

If a relevant chunk never reaches the model, I investigate retrieval. If the evidence is present but the answer misses it, I investigate context use and task design. If a completed result cannot be resumed because it was never saved, I investigate persistence. These need different checks.

Chroma’s Context Rot report found declining reliability as input length increased in its evaluated models and tasks. It supports testing how useful context is, with results tied to that evaluation. In my larger risk assessments, lower reliability remains an observation: context effects, conflicting material and task complexity are possible explanations to test. [03]

I want an agent to have enough relevant context to act, and enough durable state to explain and continue its work. A larger window alone does not establish either.

REFERENCES / 03

Sources.

This article builds on dated personal research notes. The source check covers the cited claims and links, rather than an exhaustive literature review. Document dates and any access limits appear below.

  1. Sourcegraph — Context Engineering: A Practical Guide for AI Agents (opens in a new tab)

    Engineering articlePublished Accessed

    What Is Context Engineering?; The Four Pillars; Memory.

  2. Anthropic — Effective context engineering for AI agents (opens in a new tab)

    Engineering articlePublished Accessed

    Context engineering for long-horizon tasks; Compaction; Structured note-taking.

  3. Chroma — Context Rot: How Increasing Input Tokens Impacts LLM Performance (opens in a new tab)

    Research reportPublished Accessed

    Controlled experiments across 18 evaluated LLMs; Limitations & Future Work.