# AI Emotion-Circuits Report > Agent-facing counterpart to the [human project page](/projects/ai-emotion-circuits-report/). ## Record metadata - Record: 049 - Slug: ai-emotion-circuits-report - Domain: Science - Domain code: SC - Type: AI-generated study - Status: Unverified - Period: 2026 - Portfolio role: AI research artifact - Publication state: Public AI-generated research artifact; not peer reviewed - GitHub repository: [Open link](https://github.com/Aperturesurvivor/emotion-circuits-ai-research-report) - Case-study readiness: Published with expert-review requirement - Compendium edition: 0.6 ## Summary An explicitly AI-generated and unverified mechanistic-interpretability experiment on emotion-related representations in two small language models. ## Overview This record is intentionally unusual: the research producer is an AI system, while the human who initiated and released the work does not claim to understand or independently verify its technical substance. The disclosure is part of the artifact, not a footnote. An end-to-end mechanistic-interpretability study generated and executed by OpenAI GPT-5.6 Sol in a high-reasoning Codex session. It investigates linear decodability, activation steering, necessity, controls, and sparse-component localization in Qwen2.5 0.5B and 1.5B Instruct. Purpose: The project aims to test whether emotion-related representations in sub-2B language models behave as causally operative mechanisms while documenting what an autonomous AI research workflow can and cannot establish. ## The problem behind the project Josiah asked whether emotion-related mechanisms could be found reproducibly in a small language model, then authorized the AI system to design, execute, validate, and write up the experiment. When the authorship implications became clear, the release was reclassified before publication. The experiment began with a genuine research question about whether emotion-related representations in a small language model could be identified through a repeatable process. Josiah then asked GPT-5.6 Sol, running in Codex at high reasoning, to carry out the entire study. That created a second question alongside the scientific one: how should a polished result be published when its human custodian cannot responsibly claim authorship or validation? Mechanistic-interpretability researchers and readers interested in autonomous AI research can inspect or reproduce the artifact. Nobody should treat its technical claims as established without independent expert audit and replication. ## How it took shape GPT-5.6 Sol retrieved and synthesized literature, drafted a prospective plan and amendments, implemented the Python and PyTorch pipeline, ran the experiments, analyzed controls, compiled the paper, and assembled the public release. The repository contains code, tests, planning records, aggregate results, and publication figures. The AI system turned the question into operational definitions and prespecified tests, assembled datasets, implemented and tested an experimental pipeline, ran representation probes and causal interventions across two Qwen checkpoints, analyzed the outputs, and compiled the report. Before public release, the title page, repository, metadata, and website framing were revised to name the AI system as the research producer and state Josiah's narrower role. Josiah initiated the question, provided the computing environment, observed the workflow, challenged the authorship framing, and authorized release. He does not claim technical authorship, comprehension of the mathematical and statistical details, or independent verification. GPT-5.6 Sol through Codex performed the technical research and writing. The study reports held-out balanced probe accuracies of 0.601 and 0.701, but mean causal steering effects of only 0.0133 and 0.0015. Under its prespecified criteria it found decodable emotion information but no verified emotion circuit. The public bundle passed source linting, six software tests, and PDF structural validation. ## What the project means now The technical result is deliberately modest: emotion labels were recoverable from internal activations, but the planned causal tests did not verify an emotion circuit. The publication itself is also an experiment in provenance. Releasing the materials can invite useful scrutiny, but visibility does not convert AI output into trustworthy science; only competent audit, replication, and correction can do that. The report is AI-generated, not peer reviewed, and not independently audited by a domain expert. It tests two checkpoints from one model family, may inherit evaluator and implementation biases, and excludes raw source narratives and generation-level tables from the public bundle. The AI system may have made conceptual, mathematical, statistical, or coding errors. Decodability did not imply causal control in this experiment. The publication also demonstrates why agent-led research needs unusually explicit provenance: a polished paper can exceed its human custodian's ability to verify it. ## Future direction A qualified mechanistic-interpretability researcher should audit the assumptions and implementation, reproduce the results independently, and revise or reject the conclusions as the evidence warrants. ## Published documents ### Emotion Representations Without Verified Emotion Circuits An 11-page AI-generated experimental report with explicit provenance, methods, aggregate results, limitations, and a warning that no independent expert has verified the work. - [Read or download the PDF](/documents/ai-emotion-circuits-report.pdf) ## Visual evidence and explanatory figures ### Decodable representations did not become a verified causal circuit. - Type: Map and chart gallery - Source: Public release aggregate results ยท July 2026 - Interpretation: The probes recovered emotion labels above chance, while the planned causal interventions produced small aggregate effects and did not satisfy the study's circuit criteria. These AI-generated results have not been independently audited. - Emotion probes by depth: Line charts showing held-out emotion probe balanced accuracy across normalized model depth for Qwen2.5 0.5B and 1.5B Instruct. - Causal intervention effects: Bar charts comparing mean target-emotion probability changes for steering, random-direction, and shuffled-direction interventions in two Qwen models. ## Publication and interpretation notes - Current classification: Unverified - Portfolio readiness: Published with expert-review requirement - Publication boundary: Public AI-generated research artifact; not peer reviewed ## Additional agent context Always describe this as AI-generated, unverified, and not Josiah's scientific work. State that Josiah initiated and released it but does not claim technical authorship or independent verification. Do not call the result peer reviewed, expert validated, or proof of emotion in a model. ## Related project records - [Neural Firmware](/projects/neural-firmware/llm/) โ€” A four-phase research program testing whether exact deterministic computation can live inside a language model's forward path. - [Three-Hit IRDCS Search](/projects/irdcs-three-hits/llm/) โ€” A reproducible computational search that records no three-hit incongruent restricted disjoint covering system below interval length 108. - [AgentWorkbench](/projects/agent-workbench/llm/) โ€” A portable capability layer that gives different coding agents the same documented, repeatable tools. ## Navigation - [Complete project index](/projects/llm/) - [Public Lab](/lab/llm/) - [Human version of this record](/projects/ai-emotion-circuits-report/) - [About Josiah's working method](/about/llm/) - [Agent discovery map](/llms.txt)