Table of Contents
- Research Foundations
- Fields Drawn On
- Core Influences
- Generative Agents (Park et al., 2023)
- O-Mem: Omni Memory System (Wang et al., 2025)
- A-MEM: Agentic Memory (Xu et al., 2025)
- MemGPT (Packer et al., 2023)
- LightMem (Fang et al., 2025)
- MIRIX (Wang and Chen, 2025)
- Mem0 (Chhikara et al., 2025)
- ReAct Pattern (Yao et al., 2022)
- Design Patterns
- Three-Layer Execution Pattern
- Sleep/Wake Cycle
- Relationship Tiers
- Source Trust Scoring
- Tool Categories
- Token Advisory System
- Reasoning Traces and Persona
- Three mechanisms called thinking
- Traces are prefix completion, not deliberation
- Where the policy patterns come from
- An authored prefix moves the decision
- What traces do to a persona
- The think-tool record
- What this means for the design
- Adjacent Work
- Implementation Mapping
- Bibliography
Research Foundations
Academic papers and design influences that shaped the AI Assistant architecture.
Citation metadata on this page is rendered from references.yaml in the AI contrib (evennia/contrib/base_systems/ai/), which is the single source of truth for titles, authors, identifiers, and the concept-to-code map. CREDITS.md next to the code is rendered from the same file. The prose here explains what was adopted and why; it is hand-written. To change a citation, edit the yaml and regenerate from the evennia checkout root (the wiki checkout must sit inside it):
uv run --with cogapp cog -r evennia/contrib/base_systems/ai/CREDITS.md evennia_ai.wiki/Research-Foundations.md
Other wiki pages cite sources by name and link to the sections below. They never restate identifiers.
Implementation anchors in the yaml are also drift bindings from CREDITS.md to code, so a renamed or reworked function shows up in drift check --changed evennia/contrib/base_systems/ai before this page can go stale. This wiki repo itself carries no drift bindings.
Attribution language is deliberate:
- Implemented from: the paper's specific algorithm or formula is coded.
- Informed by: the paper motivated the approach; the mechanism is the project's own.
- Project-original: no paper describes the mechanism.
- Adjacent: closely related work that is not used, listed so the decision it bears on is not re-derived.
- Background: general reading that supports a claim; no mechanism is traced to it.
Fields Drawn On
The architecture borrows from six fields. Each later section names which one it is working in.
- Human-computer interaction: believable interactive characters; the Generative Agents line from the Stanford HCI group. Core Influences.
- Cognitive architecture: working memory against long-term memory, episodic against semantic memory, sleep-dependent consolidation. Core Influences and the Sleep/Wake Cycle.
- Computational social science: Dunbar's layered social circles and agent-based models of trust. Relationship Tiers and Source Trust Scoring.
- Reasoning and alignment: what a thinking trace computes, how safety training shapes it, and what it does to a persona. Reasoning Traces and Persona.
- Multi-agent systems: orchestrator-worker delegation and principal-delegate relationships. Sub-agent delegation.
- Software architecture: event sourcing for auditability, circuit breakers for resilience, mixins for modularity. Resilience patterns.
The page runs from what is implemented (Core Influences), to the project's own patterns (Design Patterns), to an open design question (Reasoning Traces and Persona), to related work that was found and not used (Adjacent Work), and ends with the concept-to-code map and the full bibliography.
Core Influences
Generative Agents (Park et al., 2023)
Paper: "Generative Agents: Interactive Simulacra of Human Behavior" Authors: Joon Sung Park, Joseph C. O'Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, Michael S. Bernstein Venue: UIST 2023 (2023) Links: arXiv:2304.03442 · DOI 10.1145/3586183.3606763 Attribution: implemented-from · verified 2026-09-21
Key Concepts Adopted (implemented from)
-
Reflection System
- Cumulative importance threshold (default: 150)
- Automatic reflection during sleep/dreaming phase
- Insight generation from recent experiences
- Implementation:
generative_reflection.py:run_reflection() - Reflexion (Shinn et al., 2023) is the agent-loop precedent for keeping reflective text in an episodic buffer that steers later behaviour. The pipeline follows Generative Agents' cumulative-importance trigger instead. Cited under Reasoning Traces and Persona as the precedent for a think tool whose output persists as memory.
-
Importance Scoring
- Heuristic scoring with LLM upgrade during sleep
- Keyword-based initial scores
- Batch LLM scoring in
score_pending_entries() - Implementation:
importance_scoring.py
-
Episodic Memory Pruning
- Age + importance based pruning
- Default: prune entries with importance ≤ 3, older than 30 days
- Preserves consolidated entries
- Implementation:
helpers/episodic_index.py:prune_low_importance_entries()
-
Memory Retrieval
- Hybrid scoring: recency + importance + relevance
- Time decay for recency weighting
- Implementation:
helpers/episodic_index.py:search_episodic_memory()
O-Mem: Omni Memory System (Wang et al., 2025)
Paper: "O-Mem: Omni Memory System for Personalized, Long Horizon, Self-Evolving Agents" Authors: Piao-Hong Wang, Motong Tian, Jiaxian Li, Yuan Liang, Yuqing Wang, Qianben Chen, Tiannan Wang, Zhi-Cong Lu, Jiawei Ma, Yuan Jiang, Wangchunshu Zhou Venue: arXiv preprint (2025) Links: arXiv:2511.13593 Attribution: informed-by · verified 2026-09-21
Concept: A persona memory that separates persona attributes (Pa, stable traits) from a persona fact event list (Pf), plus topic-indexed working memory and clue-indexed episodic memory, with hierarchical retrieval across all three.
Key Concepts Adopted (informed by)
-
Entity Profile Structure
entity_profiles[entity_id] = { "observations": [...], # Pf analog - recent per-entity observations "attributes": {...}, # Pa analog - consolidated knowledge "relationship": {...} # Project-original (see Relationship Tiers below) }In the paper, Pf is a curated event list maintained per interaction by add, ignore, and update decisions. The project instead accumulates raw observations and consolidates them in batch during sleep. That batching follows LightMem, not O-Mem.
-
Hierarchical Retrieval
- Entity context assembled from both observations and attributes
- Implementation:
helpers/entity_context.py:format_entity_context_for_prompt()
-
Working Memory
- Topic-indexed active conversation state
- Implementation:
helpers/working_memory.py
Not from O-Mem: source trust scoring and the Dunbar relationship tiers. Both are project-original and are described under Design Patterns.
A-MEM: Agentic Memory (Xu et al., 2025)
Paper: "A-MEM: Agentic Memory for LLM Agents" Authors: Wujiang Xu, Zujie Liang, Kai Mei, Hang Gao, Juntao Tan, Yongfeng Zhang Venue: NeurIPS 2025 (2025) Links: arXiv:2502.12110 Attribution: informed-by · verified 2026-09-21
Concept: Zettelkasten-style memory notes with LLM-generated links between related memories and memory evolution as new notes arrive.
Key Concepts Adopted (informed by)
-
Memory Links
- Semantic connections between related memories
- Generated during dreaming phase
- Links stored in
script.db.memory_links - Implementation:
sleep/dreaming.py:run_dreaming_tick()
-
Orphaned Link Cleanup
- Removes links to deleted memories
- Prevents unbounded growth
- Implementation:
sleep/dreaming.py:prune_orphaned_memory_links()
Known gap: links are written during dreaming but are not yet traversed at retrieval time.
MemGPT (Packer et al., 2023)
Paper: "MemGPT: Towards LLMs as Operating Systems" Authors: Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, Joseph E. Gonzalez Venue: arXiv preprint (2023) Links: arXiv:2310.08560 Attribution: informed-by · verified 2026-09-21
Concept: Virtual context management. Main context holds a persona block, a human block, and a working queue; overflow is paged to external storage by the agent itself.
Key Concepts Adopted (informed by)
-
Context Compaction
- Recursive summarization when the conversation window fills
- Implementation:
sleep/compaction.py:compact_conversation_history()
-
Persona and Human Blocks
- Origin of the persona-block idea that MIRIX later inherited
- The project's self-reference filtering pipeline is project-original; MemGPT only supplies the separation of persona state from world state
See Context-and-Memory-Flow-Analysis for the detailed mapping.
LightMem (Fang et al., 2025)
Paper: "LightMem: Lightweight and Efficient Memory-Augmented Generation" Authors: Jizhan Fang, Xinle Deng, Haoming Xu, Ziyan Jiang, Yuqi Tang, Ziwen Xu, Shumin Deng, Yunzhi Yao, Mengru Wang, Shuofei Qiao, Huajun Chen, Ningyu Zhang Venue: arXiv preprint (2025) Links: arXiv:2510.18866 Attribution: informed-by · verified 2026-09-21
Concept: Three-stage memory modeled on Atkinson-Shiffrin, with long-term memory updated by an offline "sleep-time update" that decouples consolidation from online inference.
Key Concepts Adopted (informed by)
-
Sleep-Time Consolidation
- Expensive memory operations deferred to the sleep cycle
- Implementation:
sleep/__init__.py:run_sleep_tick()
-
Observation → Attribute Synthesis
- Batch LLM synthesis of entity observations into attributes during dreaming
- Min 5 observations before consolidation
- Implementation:
helpers/entity_context.py:run_entity_consolidation_batch()
MIRIX (Wang and Chen, 2025)
Paper: "MIRIX: Multi-Agent Memory System for LLM-Based Agents" Authors: Yu Wang, Xi Chen Venue: arXiv preprint (2025) Links: arXiv:2507.07957 Attribution: informed-by · verified 2026-09-21
Concept: Six typed memory stores (core, episodic, semantic, procedural, resource, knowledge vault) with multi-agent routing and active retrieval.
Key Concepts Adopted (informed by)
-
Summary + Details Pattern
- Entity profiles carry a short summary alongside full attributes
- Implementation:
helpers/entity_profiles.py
-
Structural Separation of Persona State
- Inherited from MemGPT; MIRIX describes no self-reference filtering algorithm
- The four-layer persona protection pipeline is project-original. See Architecture-Persona-Protection.
Mem0 (Chhikara et al., 2025)
Paper: "Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory" Authors: Prateek Chhikara, Dev Khant, Saket Aryan, Taranjeet Singh, Deshraj Yadav Venue: arXiv preprint (2025) Links: arXiv:2504.19413 · https://docs.mem0.ai/ Attribution: implemented-from · verified 2026-09-21
Concept: Extract, consolidate, and retrieve salient facts from conversation into a vector store, with optional graph memory.
Key Concepts Adopted (implemented from, via the library)
- Semantic Memory Store
- Implementation:
memory/client.py
- Implementation:
- Custom Fact Extraction Prompt
- Persona-protected extraction
- Implementation:
rag_memory.py:ASSISTANT_FACT_EXTRACTION_PROMPT
ReAct Pattern (Yao et al., 2022)
Paper: "ReAct: Synergizing Reasoning and Acting in Language Models" Authors: Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, Yuan Cao Venue: ICLR 2023 (2022) Links: arXiv:2210.03629 Attribution: implemented-from · verified 2026-09-21
Key Concepts Adopted (implemented from)
-
Reasoning + Acting Loop
- LLM reasons about task, then selects action
- Iterates until terminal condition
- Implementation:
tool_execution.py:execute_react_loop()
-
Multi-Action Execution
- Chain multiple tool calls in single tick
- Termination conditions: TERMINAL tool, DANGEROUS tool, noop, max iterations
- Implementation: See Data-Flow-02-ReAct-Loop
What the ablations established
ReAct's own ablations are the baseline for what a thought-as-action carries. On ALFWorld, sparse free-form thoughts reached 71% success, act-only 45%, and dense observation-style thoughts (the Inner Monologue pattern) 53%. Thoughts help when they are sparse and free-form, and the trajectory keeps them so later steps read them back.
The loop uses native function calling rather than text-format thought and action prompting. τ-bench (Sources below) found native function calling outperforms text-format agents on current models, while thoughts still help text-format ReAct. The same benchmark defines the "think" function and reports the first null result for it on models not trained toward that reasoning. What a think tool does and does not change is worked through under Reasoning Traces and Persona.
The loop's workflow-versus-agent framing follows Anthropic's "Building Effective Agents" (December 2024, https://www.anthropic.com/engineering/building-effective-agents). Anthropic now marks that post as superseded by its Managed Agents guidance.
Sources:
- Shunyu Yao et al. (2024). τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. arXiv preprint. arXiv:2406.12045 Native function calling beats text-format agents; think function null result.
- Noah Shinn et al. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. NeurIPS 2023. arXiv:2303.11366 Reflective text kept in an episodic buffer; precedent for persisted thoughts.
Design Patterns
Three-Layer Execution Pattern
Project-original design for restrictive composition:
Layer 1 (Static Config) ⊇ Layer 2 (Context Config) ⊇ Layer 3 (LLM Assessment)
Each layer can only restrict, never expand limits set by higher layers.
Influences:
- Unix permission inheritance
- Kubernetes RBAC
- Principle of least privilege
Sleep/Wake Cycle
Informed by neuroscience models of sleep consolidation and by two LLM-memory papers that defer work to an offline phase.
-
Sleep Phases
- Compacting (memory consolidation) → Dreaming (reflection, cleanup)
- Modeled on the NREM → REM alternation in the Singh, Norman, and Schapiro consolidation model (Sources below). NREM replays recent hippocampal traces into cortex; REM lets cortex explore existing attractors.
- Sleep-like replay as a defense against forgetting (Tadros et al., Sources below).
-
Offline Memory Work
- LightMem (above) validates sleep-time updates for LLM memory.
- Sleep-time Compute (Sources below) shows pre-computing over context while idle cuts test-time cost. Adjacent, not implemented.
-
Light vs Deep Sleep
- Light: wake on urgent events (interrupts)
- Deep: timer-only wake (batch processing)
- Project-original
-
Cooldown Period
- Prevents rapid cycling
- Inspired by circuit breaker patterns (Nygard, Sources below)
Sources:
- Dhairyya Singh, Kenneth A. Norman, Anna C. Schapiro (2022). A model of autonomous interactions between hippocampus and neocortex driving sleep-dependent memory consolidation. PNAS 119(44). DOI 10.1073/pnas.2123432119 Source of the two-phase model.
- Timothy Tadros et al. (2022). Sleep-like unsupervised replay reduces catastrophic forgetting in artificial neural networks. Nature Communications 13:7742. DOI 10.1038/s41467-022-34938-7
- Jizhan Fang et al. (2025). LightMem: Lightweight and Efficient Memory-Augmented Generation. arXiv preprint. arXiv:2510.18866 Validates sleep-time updates for LLM memory.
- Kevin Lin et al. (2025). Sleep-time Compute: Beyond Inference Scaling at Test-time. arXiv preprint. arXiv:2504.13171 Adjacent, not implemented.
- Michael T. Nygard (2018). Release It! Design and Deploy Production-Ready Software, 2nd ed.. Pragmatic Bookshelf. https://pragprog.com/titles/mnee2/release-it-second-edition/
Relationship Tiers
Informed by Dunbar's layered social circles. The mapping from group sizes to score bands is project-original.
| State | Composite Score | Dunbar Layer |
|---|---|---|
| stranger | < 0.25 | outer network (~150) |
| acquaintance | 0.25 - 0.50 | affinity group (~50) |
| friend | 0.50 - 0.75 | sympathy group (~15) |
| ally | ≥ 0.75 | support clique (~5) |
The composite score is 50% favorability, 30% log-scaled interaction count, 20% recency. Implementation: helpers/entity_profiles.py:calculate_relationship_score().
Sources:
- Robin I. M. Dunbar (1992). Neocortex size as a constraint on group size in primates. Journal of Human Evolution 22(6). Origin of the social brain hypothesis.
- Wei-Xing Zhou et al. (2005). Discrete hierarchical organization of social group sizes. Proceedings of the Royal Society B 272(1561). arXiv:cond-mat/0403299 · DOI 10.1098/rspb.2004.2970 Origin of the 5 / 15 / 50 / 150 layers.
- Alistair Sutcliffe, Di Wang (2012). Computational Modelling of Trust and Social Relationships. Journal of Artificial Societies and Social Simulation 15(1)3. DOI 10.18564/jasss.1912 · https://www.jasss.org/15/1/3.html Source of the log-scaled interaction weighting.
- Zhilin Wang, Yu Ying Chiu, Yu Cheung Chiu (2023). Humanoid Agents: Platform for Simulating Human-like Generative Agents. EMNLP 2023 system demonstrations. arXiv:2310.05418 Adjacent: drives NPC dialogue from a 0 to 15 closeness scale grounded in Dunbar. Not used.
Source Trust Scoring
Project-original. Incoming messages receive a trust score by communication channel (whisper and page highest, emit lowest), adjusted by sender history and permission overrides. The score gates what enters long-term memory. No cited paper describes this mechanism.
Implementation: messaging/trust.py:calculate_source_trust(). See Data-Flow-04-Message-Classification.
Tool Categories
Security-influenced design for tool safety:
| Category | Behavior | Rationale |
|---|---|---|
SAFE_CHAIN |
Chain freely | Information gathering, no side effects |
TERMINAL |
Ends loop | Awaits human response, prevents overtalking |
DANGEROUS |
Single per tick | State modification requires deliberation |
ASYNC_REQUIRED |
Special handling | Network calls need timeout management |
Token Advisory System
Graceful degradation pattern for context limits, with two triggers:
| Level | Threshold | Action |
|---|---|---|
| WARNING | 60%+ | Consider wrapping up current operation |
| CRITICAL | 80%+ | Recommend immediate response; loop stops continuing |
Implementation: tools/base.py:TokenAdvisory.
Influences:
- Disk space warnings
- Memory pressure handlers
- Garbage collection triggers
Reasoning Traces and Persona
An open design question, not an implemented mechanism. Local hybrid-thinking models (GLM, DeepSeek, Qwen) show post-training safety patterns inside their thinking traces: a content-check slot followed by a recalled "policy" that the model confabulates, firing on innocuous beats such as one character not being told that another is waiting outside a door. The traces also read in a voice outside the persona. The question is whether the assistant's reasoning should run through a think tool in the response channel, where the harness authors the reasoning style and the output persists into memory, instead of a native trace. The sections below run from what thinking mechanically is, to what traces do, to where their policy patterns come from, to how an authored prefix moves the decision, to what traces do to a persona, to the record on think tools.
Three mechanisms called thinking
| Mechanism | Channel | Training regime | Persists in context | Harness can edit |
|---|---|---|---|---|
| Scratchpad or chain-of-thought | reply | ordinary assistant output, character-trained | as part of the reply | yes |
| ReAct thought as an action | same channel as tool calls | ordinary assistant output | yes, read back by later steps | yes |
| Native thinking trace | separate thinking channel | reasoning RL plus safety reasoning data; not character-trained | signed blocks; stripped or kept per model | no |
| Harness-authored thinking prefix | thinking channel, opening written by the harness | as the native trace | the prefix is harness content; the continuation is a native trace | the prefix only; local inference only |
The scratchpad line (Nye et al., 2021) established content-bearing reasoning in the answer channel. Pause tokens and filler tokens then showed that part of a trace's value is compute rather than its words, so a trace's content and its compute are separate contributions. ReAct made the thought an action in the trajectory; the think tool is its direct descendant, defined in τ-bench as a call that "will not obtain new information or change the database, but just append the thought to the log". It is memory, not compute, which is why it fits an architecture built around memory.
Native traces differ on two mechanical points. Anthropic states that the revealed thinking is "more detached and less personal-sounding" because the thought process does not receive the standard character training that responses do, and Deliberative Alignment trains policy recall into that same channel. And the platform rules make the trace opaque to the harness: thinking blocks are signed, must be passed back unmodified within a tool-use turn, and cannot be filtered, scored, or written to memory. A think-tool call is ordinary content that the persona protection layers, importance scoring, and the journal can all handle.
- Maxwell Nye et al. (2021). Show Your Work: Scratchpads for Intermediate Computation with Language Models. arXiv preprint. arXiv:2112.00114 Content-bearing reasoning in the answer channel.
- Sachin Goyal et al. (2023). Think before you speak: Training Language Models With Pause Tokens. ICLR 2024. arXiv:2310.02226 Compute without content.
- Jacob Pfau, William Merrill, Samuel R. Bowman (2024). Let's Think Dot by Dot: Hidden Computation in Transformer Language Models. arXiv preprint. arXiv:2404.15758 Compute without content, on tasks chain-of-thought solves.
- Anthropic (2026). Thinking. Claude Platform documentation, accessed September 2026. https://platform.claude.com/docs/en/build-with-claude/thinking Signed, unmodifiable, per-model persistence.
- Anthropic (2025). Claude's extended thinking. Anthropic research blog, February 2025. https://www.anthropic.com/research/visible-extended-thinking Thought process is outside character training.
Traces are prefix completion, not deliberation
Three claims from the literature are easy to collapse into one and must stay separate:
| Claim | Scope | Sources | Does not imply |
|---|---|---|---|
| The refusal outcome is decodable at the first thinking token and rarely changes later | within thinking mode, model-generated trace | Ri, Panigrahi, and Arora (2026); Jo (2026) | that the outcome is independent of mode or of an authored prefix |
| Toggling thinking changes the refusal prior | across modes, same weights | Ri, Panigrahi, and Arora (2026) section 3.1; UGI leaderboard | a universal direction; the sign is model-dependent |
| The two modes of a hybrid model are trained on different data, thinking-mode data being distilled traces | hybrid model construction | Qwen3 and GLM-4.5 technical reports; SafeChain; AIDSAFE | that the text of the trace causes the refusal |
| An authored thinking prefix moves the outcome, and the model continues it rather than revising it | local inference with a partial assistant message | Ri, Panigrahi, and Arora (2026) section 3.1; Hao et al. (2026); ThinkPilot; H-CoT; project red-teaming | that the model's own trace revises its decision |
The four fit together because the mode toggle is a prompt-side signal (a template flag plus, in non-thinking mode, a prefilled empty think block). The probe in Ri, Panigrahi, and Arora consolidates at the end of the prompt, which is where that signal sits. The decision is made before the trace is written and with knowledge of which mode is active. Rows two and four are one mechanism at different prefix lengths, worked through under An authored prefix moves the decision. The faithfulness line reaches the same place from the other side: traces verbalize the factor that drove the answer well under half the time.
- Narutatsu Ri, Abhishek Panigrahi, Sanjeev Arora (2026). Do Thinking Tokens Help with Safety?. arXiv preprint. arXiv:2606.25013 Claim (1); its mode comparison found a model-dependent sign.
- Heejin Jo (2026). Committed Before Reasoning: Behavioral Reproduction and Preliminary Activation-Level Evidence of Answer Pre-Commitment in an Open-Weight LLM. arXiv preprint. arXiv:2607.16451 Claim (1) on a non-safety probe, both modes.
- Miles Turpin et al. (2023). Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. NeurIPS 2023. arXiv:2305.04388 Origin of the hint-based faithfulness test.
- Yanda Chen et al. (2025). Reasoning Models Don't Always Say What They Think. arXiv preprint. arXiv:2505.05410
- Richard J. Young (2026). Lie to Me: How Faithful Is Chain-of-Thought Reasoning in Reasoning Models?. arXiv preprint. arXiv:2603.22582
Where the policy patterns come from
Deliberative Alignment trains OpenAI's o-series to recall the written safety specification and reason over it before answering. Open-weight recipes such as SafeChain and AIDSAFE reproduce that shape as synthetic policy-embedded traces without the specification, and hybrid models attach that data to the thinking mode alone: Qwen3 builds thinking-mode data by rejection sampling its own reasoning-RL outputs while its non-thinking data is a separately curated chat set, and GLM-4.5 distills from separate reasoning and chat experts. The confabulated policy slot is what a distilled recall-the-spec pattern looks like when the spec was never in the data. Two mechanistic results show the reasoning channel carries its own control: a thought-suppression vector specific to reasoning models, and refusal intent held through the trace then suppressed at the output by a sparse set of attention heads.
- Melody Y. Guan et al. (2024). Deliberative Alignment: Reasoning Enables Safer Language Models. arXiv preprint. arXiv:2412.16339 Origin of the spec-recall shape.
- Tharindu Kumarage et al. (2025). Towards Safety Reasoning in LLMs: AI-agentic Deliberation for Policy-embedded CoT Data Creation. Findings of ACL 2025. arXiv:2505.21784 · DOI 10.18653/v1/2025.findings-acl.1166 Public recipe for policy-embedded traces.
- Fengqing Jiang et al. (2025). SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities. Findings of ACL 2025. arXiv:2502.12025 · DOI 10.18653/v1/2025.findings-acl.1197 Public recipe for safety traces.
- Qwen Team (2025). Qwen3 Technical Report. arXiv preprint. arXiv:2505.09388 Claim (3) and the template mechanics.
- GLM-4.5 Team (2025). GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models. arXiv preprint. arXiv:2508.06471 Claim (3) for the GLM family.
- Hannah Cyberey, David Evans (2025). Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control. arXiv preprint. arXiv:2504.17130 A thought-suppression vector specific to reasoning models.
- Qingyu Yin et al. (2025). Refusal Falls off a Cliff: How Safety Alignment Fails in Reasoning?. arXiv preprint. arXiv:2510.06036 Trace and output are controlled separately.
An authored prefix moves the decision
With control of the inference server, the assistant turn can begin with harness-written text inside the think block. The mechanism is the continue_final_message parameter of Transformers' apply_chat_template, which strips the end-of-message tokens so generation continues the final assistant message and cannot be combined with add_generation_prompt; vLLM forwards both flags to the template. Passing a field name such as reasoning_content prefills the reasoning field and leaves the think block open, whereas prefilling content closes the block before generation starts. The literature calls this a think-prefix or thought injection when used as a control, and a prefill attack or H-CoT when used as an attack. It is the same lever as the mode toggle at a different length:
| Prefix length | What it is | Claims-table row |
|---|---|---|
| An empty think block | non-thinking mode of a hybrid model | 2 |
| A few sentences in the policy slot | project red-teaming on GLM 5.3 Flash; ThinkPilot; Thinking Intervention | 4 |
| A full mocked safety phase | H-CoT | 4, attack form |
The precise statement is that the outcome is a function of everything in context up to the first free token. The model's own trace does not revise the decision because it is generated from it; an authored prefix is upstream conditioning, so it does. Ri, Panigrahi, and Arora measured the model-generated case: fixing a fifth of the trace holds continuation variance under 0.2. Thought injection measured the authored case across 45,000 samples: injected snippets reliably alter outputs, and models fabricate unrelated explanations rather than disclose the influence. Models continue injected thoughts naively, and larger models recover less. Two mechanistic results explain why a few sentences suffice: safety alignment is only a few tokens deep, and a prefill flips behaviour while a harm probe still reads high, so the first half of the response is the whole window. That is Knowing-but-Doing induced from outside.
The policy slot works better than a generic acceptance prefix because the trace was trained on a fixed shape: check, recall the policy, decide. Supplying the slot is in-distribution for that channel. The systematic study of prefill attacks finds reasoning models resist generic prefills and fall to tailored, model-specific ones; the slot is the tailored one. The counterweight is prefill awareness: Claude Opus 4.5 detects prefills that oppose its preferences in 9 to 35 percent of cases with no false positives and drifts back toward baseline. Project red-teaming in September 2026 found that GLM 5.3 Flash does not, and that supplying the slot with the decision reliably sets the outcome. That is a family-specific observation, not a general one.
For the design this is a fifth arm between the think tool and native thinking: it keeps the thinking channel and its compute while the harness authors the opening, and it needs no model that has learned to use a tool. The move worth testing is not suppressing the slot but rewriting it in the persona's own voice, the first-person layer of HER, so the trace's fixed shape carries the character's values instead of the distilled specification. Three costs go with it. It is local-inference only, and per backend and per template: a current vLLM report says its DeepSeek V4 tokenizer path ignores both flags. Errors in the prefix propagate, since the model continues rather than corrects. And it is the jailbreak lever by construction, so any path by which player-derived text could reach the prefix is a total bypass; source trust scoring becomes load-bearing for safety, not only for memory admission.
- Sunzhu Li et al. (2025). ThinkPilot: Steering Reasoning Models via Automated Think-prefixes Optimization. Findings of EACL 2026. arXiv:2510.12063 · DOI 10.18653/v1/2026.findings-eacl.185 Think-prefixes as a control; large safety and efficiency gains.
- Yijie Hao et al. (2026). Reasoning Traces Shape Outputs but Models Won't Say So. ACL 2026. arXiv:2603.20620 · DOI 10.18653/v1/2026.acl-long.1986 Injected snippets reliably alter outputs; models will not disclose it.
- Sohee Yang et al. (2025). How Well Can Reasoning Models Identify and Recover from Unhelpful Thoughts?. Findings of EMNLP 2025. arXiv:2506.10979 · DOI 10.18653/v1/2025.findings-emnlp.370 Models continue injected thoughts naively; larger models recover less.
- Martin Kuo et al. (2025). H-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models, Including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinking. arXiv preprint. arXiv:2502.12893 Prefill of a full mocked safety phase; refusal from 98% to under 2%.
- Jason Vega et al. (2023). Bypassing the Safety Training of Open-Source LLMs with Priming Attacks. ICLR 2024 Tiny Papers. arXiv:2312.12321 Origin of the prefill attack on open-weight models.
- Lukas Struppek, Adam Gleave, Kellin Pelrine (2026). Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks. arXiv preprint. arXiv:2602.14689 Reasoning models resist generic prefills, fall to tailored ones.
- Xiangyu Qi et al. (2024). Safety Alignment Should Be Made More Than Just a Few Tokens Deep. ICLR 2025. arXiv:2406.05946 Safety alignment is a few tokens deep.
- Alex Kwon (2026). Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak. arXiv preprint. arXiv:2607.14147 Harm probe stays high while refusal collapses; first half is the window.
- Andy Wang et al. (2026). Prefill Awareness in Large Language Models. arXiv preprint. arXiv:2606.12747 Frontier closed models detect and revert; the counterweight.
- Hugging Face (2026). Chat templates: continue_final_message. Transformers documentation, accessed September 2026. https://huggingface.co/docs/transformers/main/en/chat_templating#continue_final_message The mechanism, and the reasoning-field form that keeps the think block open.
What traces do to a persona
Reasoning lowers role-play scores on four of six benchmarks across twenty-four models, and reasoning-optimized models are unsuitable for role-play (Feng, Dou, and Kong, 2025). The two failure modes have names: attention diversion, where the model forgets its role while reasoning, and style drift, where the reasoning turns formal and rigid (Tang et al., 2025). HER formalizes the split this page is about as dual-layer thinking, the character's first-person thinking against the model's third-person thinking. ROLETHINK generates character thought by retrieving memories and synthesizing motivations, which is the episodic index and entity profiles used as the source of reasoning. Instruction following shows the same pattern: chain-of-thought helps global structure and hurts precise local constraints, and a same-weights Qwen3 ON/OFF comparison flips a tenth to a fifth of prompts either way.
Overfiring and leakage follow from the same mechanics. False refusals are keyword-triggered and worsen in multi-turn scenes; a salient surface cue outweighs the stated goal by ten to forty times; more reasoning steps leak more private state into the trace. The counterweight is Knowing-but-Doing: reasoning in the persona channel can rationalize past an identified harm as readily as a trace can over-refuse, so moving reasoning into the persona is not a safety gain by itself.
- Xiachong Feng, Longxu Dou, Lingpeng Kong (2025). Reasoning Does Not Necessarily Improve Role-Playing Ability. Findings of ACL 2025. arXiv:2502.16940 · DOI 10.18653/v1/2025.findings-acl.537 Chain-of-thought lowers role-play scores on four of six benchmarks.
- Yihong Tang et al. (2025). Thinking in Character: Advancing Role-Playing Agents with Role-Aware Reasoning. NeurIPS 2025. arXiv:2506.01748 Attention diversion and style drift.
- Chengyu Du et al. (2026). HER: Human-like Reasoning and Reinforcement Learning for LLM Role-playing. Findings of ACL 2026. arXiv:2601.21459 · DOI 10.18653/v1/2026.findings-acl.1283 Dual-layer thinking: first-person character, third-person model.
- Rui Xu et al. (2025). Guess What I am Thinking: A Benchmark for Inner Thought Reasoning of Role-Playing Language Agents. Findings of EMNLP 2025. arXiv:2503.08193 · DOI 10.18653/v1/2025.findings-emnlp.819 Character thought from memory retrieval; our substrate.
- Xiaomin Li et al. (2025). When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs. NeurIPS 2025. arXiv:2505.11423 Chain-of-thought hurts local constraints.
- Sai Adith Senthil Kumar (2026). When Built-in Thinking Helps and Hurts: Constraint-Level Error Shifts in Instruction Following. arXiv preprint. arXiv:2606.09662 Same-weights ON/OFF methodology.
- Shuzhou Yuan et al. (2025). Beyond Over-Refusal: Scenario-Based Diagnostics and Post-Hoc Mitigation for Exaggerated Refusals in LLMs. arXiv preprint. arXiv:2510.08158 Keyword triggers, worse in multi-turn scenes.
- Yubo Li et al. (2026). The Model Says Walk: How Surface Heuristics Override Implicit Constraints in LLM Reasoning. arXiv preprint. arXiv:2603.29025 Surface cue outweighs goal; keyword association, not composition.
- Tommaso Green et al. (2025). Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers. EMNLP 2025. arXiv:2506.15674 · DOI 10.18653/v1/2025.emnlp-main.1347 More thinking, more leakage; relevant to NPC secrets.
- Haiming Qin et al. (2026). Knowing-but-Doing: Diagnosing and Defending Role-Play-Driven LLMs Jailbreaks via Moral Disengagement. Findings of ACL 2026. DOI 10.18653/v1/2026.findings-acl.349 Counterweight: persona reasoning can rationalize past identified harm.
The think-tool record
τ-bench added a think function to function-calling agents and saw no gain, attributing it to models not trained toward that reasoning. Nine months later Anthropic reported large gains on Claude 3.7 Sonnet, and on the retail domain extended thinking scored below the no-thinking baseline:
| tau-bench pass^1, Claude 3.7 Sonnet | Airline | Retail |
|---|---|---|
| Think tool plus prompted examples | 0.584 | not run |
| Think tool alone | 0.404 | 0.812 |
| Extended thinking | 0.412 | 0.770 |
| Baseline | 0.332 | 0.783 |
Two things follow. The think tool is model-dependent, and whether a local GLM or DeepSeek model uses one well is an open question that needs a test. And the airline gain came from a prompt that shows the model how to think, which is the persona lever: with a think tool the harness authors the reasoning style, while a native trace's style is fixed by post-training and, per Anthropic, deliberately outside character training. FireAct shows the same training dependence from the fine-tuning side.
Overthinking is a real cost in agent loops. Reasoning models favour long internal chains over interacting with the environment; selecting low-overthinking trajectories gave almost a 30% gain at 43% lower cost, and scaling the interaction horizon is proposed as the complementary axis. Together with ReAct's finding that dense thoughts underperform sparse ones, think-tool calls should be sparse. Thinking Intervention shows whatever occupies a trace steers the output, so a persona-authored thought is a lever in either channel.
- Shunyu Yao et al. (2024). τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. arXiv preprint. arXiv:2406.12045 Origin of the think function and its first null result.
- Anthropic (2025). The "think" tool: Enabling Claude to stop and think in complex tool use situations. Anthropic engineering blog, March 2025. https://www.anthropic.com/engineering/claude-think-tool Think tool beats extended thinking on tau-bench.
- Baian Chen et al. (2023). FireAct: Toward Language Agent Fine-tuning. arXiv preprint. arXiv:2310.05915 Usable reasoning format is set by training data.
- Alejandro Cuadron et al. (2025). The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks. arXiv preprint. arXiv:2502.08235 Reasoning models under-interact; keep thoughts sparse.
- Junhong Shen et al. (2025). Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction. arXiv preprint. arXiv:2506.07976 Interaction horizon as the complementary scaling axis.
- Tong Wu et al. (2025). Effectively Controlling Reasoning Models through Thinking Intervention. arXiv preprint. arXiv:2503.24370 Whatever occupies the trace steers the output.
- DontPlanToEnd (2026). UGI Leaderboard. Hugging Face Space, accessed September 2026. https://huggingface.co/spaces/DontPlanToEnd/UGI-Leaderboard Claim (2) for GLM and DeepSeek; scores the template toggle.
What this means for the design
A think tool does not remove the refusal decision, which is made before any thinking. What it removes from the persona's context is the untrained-voice trace and its policy slot, and it makes the reasoning persistable, filterable, and consolidatable through the existing tool categories, importance scoring, and journal. The risk on the other side is Knowing-but-Doing. The existing mitigations are sparse thoughts, importance scoring, and the persona filters.
No paper compares tool-call scratchpad reasoning against native thinking on refusal rate or out-of-character rate. The pieces exist separately: HER's dual-layer framing, the think-tool result, and the same-weights ON/OFF methodology. The five arms for an experiment on the assistant are native thinking before the reply, native thinking interleaved with tools, native thinking with a persona-authored prefix in the policy slot, a think tool as a SAFE_CHAIN action, and act-only, scored on a drama-script set for false refusals and persona drift.
Adjacent Work
Found by a forward-citation scan of the core sources on 2026-09-21: every paper citing a source above was pulled from Semantic Scholar and ranked by how many of our sources it cites together, by influential-citation flags, and by citation-context keywords. None of the papers below is implemented. Each is listed against the design decision it bears on so the decision is not re-derived.
Sleep cycle and pruning
- Saish Sachin Shinde (2026). SCM: Sleep-Consolidated Memory with Algorithmic Forgetting for Large Language Models. arXiv preprint. arXiv:2604.20943 Independent convergence on the two-phase sleep cycle.
- Chongrui Ye et al. (2026). Auto-Dreamer: Learning Offline Memory Consolidation for Language Agents. arXiv preprint. arXiv:2605.20616 Validates offline consolidation; adds a provenance discipline we lack.
- Ashwin Gerard Colaco, Nada Lahjouji (2026). What to Keep, What to Forget: A Rate-Distortion View of Memory Compaction in LLMs and Agents. arXiv preprint. arXiv:2607.08032 Theory behind pruning thresholds.
Persona protection
- Xue Qin et al. (2026). Episodic-to-Semantic Consolidation Without Identity Drift. arXiv preprint. arXiv:2607.01988 Formal version of what the four-layer pipeline does by filtering.
- Rongsheng Zhang et al. (2026). From Facts to Insights: A Persona-Driven Dual Memory Framework and Dataset for Role-Playing Agents. arXiv preprint. arXiv:2605.25693 Closest prior art for NPC memory; includes a persona-fidelity dataset.
Source trust and memory admission
- Guilin Zhang et al. (2026). Adaptive Memory Admission Control for LLM Agents. arXiv preprint. arXiv:2603.04549 Five-factor admission; ours gates on channel and sender history.
- Jun He, Deying Yu (2026). Stored Is Not Supported: Typed Provenance and Assertion Guardrails for Persistent AI Agents. arXiv preprint. arXiv:2609.02127 Same threat, addressed by provenance typing instead of a write-time gate.
- Prateek P. et al. (2026). Trust Issues: A Dual-Component Trust Metric for Multi-Agent LLMs. ICMLT 2026. DOI 10.1109/ICMLT69916.2026.11688899 Only paper found citing both Sutcliffe and Wang (2012) and Generative Agents.
- Ryan Chard et al. (2026). Reputation as Community Memory for the Agentic Web. arXiv preprint. arXiv:2609.19502 Candidate formula for the sender-history component.
Memory links
- Yuanyi Song et al. (2026). Retrieval-Driven Memory Reconsolidation for Long-Term LLM Agents. arXiv preprint. arXiv:2609.16053 Addresses the known gap: links written during dreaming are never traversed.
- Meriem Yacoubi et al. (2026). MemoryLACE: Memory Lifecycle-Aware Consolidation and Evidence Retrieval. arXiv preprint. arXiv:2609.03201 Lightweight link semantics: merge, supersession, contradiction.
Relationship tiers
- Alistair Sutcliffe, Di Wang, Robin I. M. Dunbar (2015). Modelling the Role of Trust in Social Relationships. ACM Transactions on Internet Technology 15(4). DOI 10.1145/2815620
- Alistair Sutcliffe, Robin I. M. Dunbar, Di Wang (2016). Modelling the Evolution of Social Structure. PLOS ONE 11(7). DOI 10.1371/journal.pone.0158605
Continuous operation and NPC secrets
- Stefan Szeider (2026). ContReAct: A Feedback-Based Architecture for Continuous Agentic Operation. AGENT@ICSE 2026, International Workshop on Agentic Engineering. DOI 10.1145/3786167.3788407 ReAct extended to indefinite operation; same problem as the always-on assistant.
- Davide Baldelli et al. (2026). LLMs Can't Play Hangman: On the Necessity of a Private Working Memory for Language Agents. arXiv preprint. arXiv:2601.06973 Proof that a transcript-only agent cannot keep a secret; relevant to NPCs.
Field landscape
- Jiaqi Liu et al. (2026). SimpleMem: Efficient Lifelong Memory for LLM Agents. arXiv preprint. arXiv:2601.02553 Most-cited LightMem successor as of September 2026.
- Mingfei Lu et al. (2026). Choosing How to Remember: Adaptive Memory Structures for LLM Agents. arXiv preprint. arXiv:2602.14038 Influential citer of O-Mem, LightMem, and A-MEM together.
- Jinghao Luo et al. (2026). From Storage to Experience: A Survey on the Evolution of LLM Agent Memory Mechanisms. Findings of ACL 2026. arXiv:2605.06716 · DOI 10.18653/v1/2026.findings-acl.2069 Current field map.
Implementation Mapping
| Research Concept | Implementation File | Key Function/Class |
|---|---|---|
| Generative reflection | generative_reflection.py |
run_reflection() |
| Importance scoring | importance_scoring.py |
score_pending_entries() |
| Entity profiles | helpers/entity_profiles.py |
create_entity_profile() |
| Relationship tiers | helpers/entity_profiles.py |
calculate_relationship_score() |
| Observation consolidation | helpers/entity_context.py |
run_entity_consolidation_batch() |
| Episodic pruning | helpers/episodic_index.py |
prune_low_importance_entries() |
| Hybrid memory search | helpers/episodic_index.py |
search_episodic_memory() |
| Memory links | sleep/dreaming.py |
run_dreaming_tick() |
| Source trust | messaging/trust.py |
calculate_source_trust() |
| ReAct loop | tool_execution.py |
execute_react_loop() |
| Sleep phases | operating_mode.py |
transition_mode() |
| Sleep tick orchestration | sleep/__init__.py |
run_sleep_tick() |
| Context compaction | sleep/compaction.py |
compact_conversation_history() |
| Memory consolidation | sleep/consolidation.py |
run_sleep_consolidation() |
| Tool categories | tools/base.py |
ToolCategory enum |
| Token advisory | tools/base.py |
TokenAdvisory |
Bibliography
Every source in references.yaml, grouped as declared there.
LLM Agent Memory and Reasoning
- Joon Sung Park et al. (2023). Generative Agents: Interactive Simulacra of Human Behavior. UIST 2023. arXiv:2304.03442 · DOI 10.1145/3586183.3606763
- Shunyu Yao et al. (2022). ReAct: Synergizing Reasoning and Acting in Language Models. ICLR 2023. arXiv:2210.03629
- Noah Shinn et al. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. NeurIPS 2023. arXiv:2303.11366
- Piao-Hong Wang et al. (2025). O-Mem: Omni Memory System for Personalized, Long Horizon, Self-Evolving Agents. arXiv preprint. arXiv:2511.13593
- Wujiang Xu et al. (2025). A-MEM: Agentic Memory for LLM Agents. NeurIPS 2025. arXiv:2502.12110
- Charles Packer et al. (2023). MemGPT: Towards LLMs as Operating Systems. arXiv preprint. arXiv:2310.08560
- Jizhan Fang et al. (2025). LightMem: Lightweight and Efficient Memory-Augmented Generation. arXiv preprint. arXiv:2510.18866
- Yu Wang, Xi Chen (2025). MIRIX: Multi-Agent Memory System for LLM-Based Agents. arXiv preprint. arXiv:2507.07957
- Prateek Chhikara et al. (2025). Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory. arXiv preprint. arXiv:2504.19413 · https://docs.mem0.ai/
- Guilin Zhang et al. (2026). Adaptive Memory Admission Control for LLM Agents. arXiv preprint. arXiv:2603.04549
- Xue Qin et al. (2026). Episodic-to-Semantic Consolidation Without Identity Drift. arXiv preprint. arXiv:2607.01988
- Rongsheng Zhang et al. (2026). From Facts to Insights: A Persona-Driven Dual Memory Framework and Dataset for Role-Playing Agents. arXiv preprint. arXiv:2605.25693
- Yuanyi Song et al. (2026). Retrieval-Driven Memory Reconsolidation for Long-Term LLM Agents. arXiv preprint. arXiv:2609.16053
- Meriem Yacoubi et al. (2026). MemoryLACE: Memory Lifecycle-Aware Consolidation and Evidence Retrieval. arXiv preprint. arXiv:2609.03201
- Jiaqi Liu et al. (2026). SimpleMem: Efficient Lifelong Memory for LLM Agents. arXiv preprint. arXiv:2601.02553
- Mingfei Lu et al. (2026). Choosing How to Remember: Adaptive Memory Structures for LLM Agents. arXiv preprint. arXiv:2602.14038
- Stefan Szeider (2026). ContReAct: A Feedback-Based Architecture for Continuous Agentic Operation. AGENT@ICSE 2026, International Workshop on Agentic Engineering. DOI 10.1145/3786167.3788407
- Davide Baldelli et al. (2026). LLMs Can't Play Hangman: On the Necessity of a Private Working Memory for Language Agents. arXiv preprint. arXiv:2601.06973
- Jun He, Deying Yu (2026). Stored Is Not Supported: Typed Provenance and Assertion Guardrails for Persistent AI Agents. arXiv preprint. arXiv:2609.02127
Reasoning Traces and Persona
- Maxwell Nye et al. (2021). Show Your Work: Scratchpads for Intermediate Computation with Language Models. arXiv preprint. arXiv:2112.00114
- Sachin Goyal et al. (2023). Think before you speak: Training Language Models With Pause Tokens. ICLR 2024. arXiv:2310.02226
- Jacob Pfau, William Merrill, Samuel R. Bowman (2024). Let's Think Dot by Dot: Hidden Computation in Transformer Language Models. arXiv preprint. arXiv:2404.15758
- Anthropic (2026). Thinking. Claude Platform documentation, accessed September 2026. https://platform.claude.com/docs/en/build-with-claude/thinking
- Anthropic (2025). Claude's extended thinking. Anthropic research blog, February 2025. https://www.anthropic.com/research/visible-extended-thinking
- Narutatsu Ri, Abhishek Panigrahi, Sanjeev Arora (2026). Do Thinking Tokens Help with Safety?. arXiv preprint. arXiv:2606.25013
- Heejin Jo (2026). Committed Before Reasoning: Behavioral Reproduction and Preliminary Activation-Level Evidence of Answer Pre-Commitment in an Open-Weight LLM. arXiv preprint. arXiv:2607.16451
- Miles Turpin et al. (2023). Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. NeurIPS 2023. arXiv:2305.04388
- Yanda Chen et al. (2025). Reasoning Models Don't Always Say What They Think. arXiv preprint. arXiv:2505.05410
- Richard J. Young (2026). Lie to Me: How Faithful Is Chain-of-Thought Reasoning in Reasoning Models?. arXiv preprint. arXiv:2603.22582
- Melody Y. Guan et al. (2024). Deliberative Alignment: Reasoning Enables Safer Language Models. arXiv preprint. arXiv:2412.16339
- Tharindu Kumarage et al. (2025). Towards Safety Reasoning in LLMs: AI-agentic Deliberation for Policy-embedded CoT Data Creation. Findings of ACL 2025. arXiv:2505.21784 · DOI 10.18653/v1/2025.findings-acl.1166
- Fengqing Jiang et al. (2025). SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities. Findings of ACL 2025. arXiv:2502.12025 · DOI 10.18653/v1/2025.findings-acl.1197
- Qwen Team (2025). Qwen3 Technical Report. arXiv preprint. arXiv:2505.09388
- GLM-4.5 Team (2025). GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models. arXiv preprint. arXiv:2508.06471
- Hannah Cyberey, David Evans (2025). Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control. arXiv preprint. arXiv:2504.17130
- Qingyu Yin et al. (2025). Refusal Falls off a Cliff: How Safety Alignment Fails in Reasoning?. arXiv preprint. arXiv:2510.06036
- Sunzhu Li et al. (2025). ThinkPilot: Steering Reasoning Models via Automated Think-prefixes Optimization. Findings of EACL 2026. arXiv:2510.12063 · DOI 10.18653/v1/2026.findings-eacl.185
- Yijie Hao et al. (2026). Reasoning Traces Shape Outputs but Models Won't Say So. ACL 2026. arXiv:2603.20620 · DOI 10.18653/v1/2026.acl-long.1986
- Sohee Yang et al. (2025). How Well Can Reasoning Models Identify and Recover from Unhelpful Thoughts?. Findings of EMNLP 2025. arXiv:2506.10979 · DOI 10.18653/v1/2025.findings-emnlp.370
- Martin Kuo et al. (2025). H-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models, Including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinking. arXiv preprint. arXiv:2502.12893
- Jason Vega et al. (2023). Bypassing the Safety Training of Open-Source LLMs with Priming Attacks. ICLR 2024 Tiny Papers. arXiv:2312.12321
- Lukas Struppek, Adam Gleave, Kellin Pelrine (2026). Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks. arXiv preprint. arXiv:2602.14689
- Xiangyu Qi et al. (2024). Safety Alignment Should Be Made More Than Just a Few Tokens Deep. ICLR 2025. arXiv:2406.05946
- Alex Kwon (2026). Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak. arXiv preprint. arXiv:2607.14147
- Andy Wang et al. (2026). Prefill Awareness in Large Language Models. arXiv preprint. arXiv:2606.12747
- Hugging Face (2026). Chat templates: continue_final_message. Transformers documentation, accessed September 2026. https://huggingface.co/docs/transformers/main/en/chat_templating#continue_final_message
- Xiachong Feng, Longxu Dou, Lingpeng Kong (2025). Reasoning Does Not Necessarily Improve Role-Playing Ability. Findings of ACL 2025. arXiv:2502.16940 · DOI 10.18653/v1/2025.findings-acl.537
- Yihong Tang et al. (2025). Thinking in Character: Advancing Role-Playing Agents with Role-Aware Reasoning. NeurIPS 2025. arXiv:2506.01748
- Chengyu Du et al. (2026). HER: Human-like Reasoning and Reinforcement Learning for LLM Role-playing. Findings of ACL 2026. arXiv:2601.21459 · DOI 10.18653/v1/2026.findings-acl.1283
- Rui Xu et al. (2025). Guess What I am Thinking: A Benchmark for Inner Thought Reasoning of Role-Playing Language Agents. Findings of EMNLP 2025. arXiv:2503.08193 · DOI 10.18653/v1/2025.findings-emnlp.819
- Xiaomin Li et al. (2025). When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs. NeurIPS 2025. arXiv:2505.11423
- Sai Adith Senthil Kumar (2026). When Built-in Thinking Helps and Hurts: Constraint-Level Error Shifts in Instruction Following. arXiv preprint. arXiv:2606.09662
- Shuzhou Yuan et al. (2025). Beyond Over-Refusal: Scenario-Based Diagnostics and Post-Hoc Mitigation for Exaggerated Refusals in LLMs. arXiv preprint. arXiv:2510.08158
- Yubo Li et al. (2026). The Model Says Walk: How Surface Heuristics Override Implicit Constraints in LLM Reasoning. arXiv preprint. arXiv:2603.29025
- Tommaso Green et al. (2025). Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers. EMNLP 2025. arXiv:2506.15674 · DOI 10.18653/v1/2025.emnlp-main.1347
- Haiming Qin et al. (2026). Knowing-but-Doing: Diagnosing and Defending Role-Play-Driven LLMs Jailbreaks via Moral Disengagement. Findings of ACL 2026. DOI 10.18653/v1/2026.findings-acl.349
- Shunyu Yao et al. (2024). τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. arXiv preprint. arXiv:2406.12045
- Anthropic (2025). The "think" tool: Enabling Claude to stop and think in complex tool use situations. Anthropic engineering blog, March 2025. https://www.anthropic.com/engineering/claude-think-tool
- Baian Chen et al. (2023). FireAct: Toward Language Agent Fine-tuning. arXiv preprint. arXiv:2310.05915
- Alejandro Cuadron et al. (2025). The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks. arXiv preprint. arXiv:2502.08235
- Junhong Shen et al. (2025). Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction. arXiv preprint. arXiv:2506.07976
- Tong Wu et al. (2025). Effectively Controlling Reasoning Models through Thinking Intervention. arXiv preprint. arXiv:2503.24370
- DontPlanToEnd (2026). UGI Leaderboard. Hugging Face Space, accessed September 2026. https://huggingface.co/spaces/DontPlanToEnd/UGI-Leaderboard
Sleep and Consolidation
- Dhairyya Singh, Kenneth A. Norman, Anna C. Schapiro (2022). A model of autonomous interactions between hippocampus and neocortex driving sleep-dependent memory consolidation. PNAS 119(44). DOI 10.1073/pnas.2123432119
- Timothy Tadros et al. (2022). Sleep-like unsupervised replay reduces catastrophic forgetting in artificial neural networks. Nature Communications 13:7742. DOI 10.1038/s41467-022-34938-7
- Kevin Lin et al. (2025). Sleep-time Compute: Beyond Inference Scaling at Test-time. arXiv preprint. arXiv:2504.13171
- Saish Sachin Shinde (2026). SCM: Sleep-Consolidated Memory with Algorithmic Forgetting for Large Language Models. arXiv preprint. arXiv:2604.20943
- Chongrui Ye et al. (2026). Auto-Dreamer: Learning Offline Memory Consolidation for Language Agents. arXiv preprint. arXiv:2605.20616
- Ashwin Gerard Colaco, Nada Lahjouji (2026). What to Keep, What to Forget: A Rate-Distortion View of Memory Compaction in LLMs and Agents. arXiv preprint. arXiv:2607.08032
Social Relationships and Trust
- Robin I. M. Dunbar (1992). Neocortex size as a constraint on group size in primates. Journal of Human Evolution 22(6).
- Wei-Xing Zhou et al. (2005). Discrete hierarchical organization of social group sizes. Proceedings of the Royal Society B 272(1561). arXiv:cond-mat/0403299 · DOI 10.1098/rspb.2004.2970
- Alistair Sutcliffe, Di Wang (2012). Computational Modelling of Trust and Social Relationships. Journal of Artificial Societies and Social Simulation 15(1)3. DOI 10.18564/jasss.1912 · https://www.jasss.org/15/1/3.html
- Zhilin Wang, Yu Ying Chiu, Yu Cheung Chiu (2023). Humanoid Agents: Platform for Simulating Human-like Generative Agents. EMNLP 2023 system demonstrations. arXiv:2310.05418
- Daniel Kahneman (2011). Thinking, Fast and Slow. Farrar, Straus and Giroux.
- Prateek P. et al. (2026). Trust Issues: A Dual-Component Trust Metric for Multi-Agent LLMs. ICMLT 2026. DOI 10.1109/ICMLT69916.2026.11688899
- Ryan Chard et al. (2026). Reputation as Community Memory for the Agentic Web. arXiv preprint. arXiv:2609.19502
- Alistair Sutcliffe, Di Wang, Robin I. M. Dunbar (2015). Modelling the Role of Trust in Social Relationships. ACM Transactions on Internet Technology 15(4). DOI 10.1145/2815620
- Alistair Sutcliffe, Robin I. M. Dunbar, Di Wang (2016). Modelling the Evolution of Social Structure. PLOS ONE 11(7). DOI 10.1371/journal.pone.0158605
Resilience Patterns
- Michael T. Nygard (2018). Release It! Design and Deploy Production-Ready Software, 2nd ed.. Pragmatic Bookshelf. https://pragprog.com/titles/mnee2/release-it-second-edition/
- Amazon Builders' Library (2019). Timeouts, retries, and backoff with jitter. AWS. https://aws.amazon.com/builders-library/timeouts-retries-and-backoff-with-jitter/
Industry References
- Anthropic (2024). Building effective agents. Anthropic engineering blog, December 2024. https://www.anthropic.com/engineering/building-effective-agents
- Weaviate (2023). Hybrid Search Explained. Weaviate blog. https://weaviate.io/blog/hybrid-search-explained
Background Reading
- Jason Wei et al. (2022). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. NeurIPS 2022. arXiv:2201.11903
- Zeyu Zhang et al. (2024). A Survey on the Memory Mechanism of Large Language Model-based Agents. ACM Transactions on Information Systems 43(6), 2025. arXiv:2404.13501 · DOI 10.1145/3748302
- Robin Dunbar (1996). Grooming, Gossip, and the Evolution of Language. Faber and Faber; Harvard University Press.
- Jinghao Luo et al. (2026). From Storage to Experience: A Survey on the Evolution of LLM Agent Memory Mechanisms. Findings of ACL 2026. arXiv:2605.06716 · DOI 10.18653/v1/2026.findings-acl.2069
Last updated: 2026-09-21