This record is intentionally unusual: the research producer is an AI system, while the human who initiated and released the work does not claim to understand or independently verify its technical substance. The disclosure is part of the artifact, not a footnote.
The project takes the form of an end-to-end mechanistic-interpretability study generated and executed by OpenAI GPT-5.6 Sol in a high-reasoning Codex session. It investigates linear decodability, activation steering, necessity, controls, and sparse-component localization in Qwen2.5 0.5B and 1.5B Instruct. Its purpose is to test whether emotion-related representations in sub-2B language models behave as causally operative mechanisms while documenting what an autonomous AI research workflow can and cannot establish.
Aggregate results · unverified AI experiment
Decodable representations did not become a verified causal circuit.
Emotion probes by depth
Causal intervention effects
The probes recovered emotion labels above chance, while the planned causal interventions produced small aggregate effects and did not satisfy the study's circuit criteria. These AI-generated results have not been independently audited.
Source: Public release aggregate results · July 2026The question behind AI Emotion-Circuits Report
The experiment began with a genuine research question about whether emotion-related representations in a small language model could be identified through a repeatable process. Josiah then asked GPT-5.6 Sol, running in Codex at high reasoning, to carry out the entire study. That created a second question alongside the scientific one: how should a polished result be published when its human custodian cannot responsibly claim authorship or validation? Mechanistic-interpretability researchers and readers interested in autonomous AI research can inspect or reproduce the artifact. Nobody should treat its technical claims as established without independent expert audit and replication.
How AI Emotion-Circuits Report took shape
The AI system turned the question into operational definitions and prespecified tests, assembled datasets, implemented and tested an experimental pipeline, ran representation probes and causal interventions across two Qwen checkpoints, analyzed the outputs, and compiled the report. Before public release, the title page, repository, metadata, and website framing were revised to name the AI system as the research producer and state Josiah's narrower role. Josiah initiated the question, provided the computing environment, observed the workflow, challenged the authorship framing, and authorized release. He does not claim technical authorship, comprehension of the mathematical and statistical details, or independent verification. GPT-5.6 Sol through Codex performed the technical research and writing.
What the evidence supports
The study reports held-out balanced probe accuracies of 0.601 and 0.701, but mean causal steering effects of only 0.0133 and 0.0015. Under its prespecified criteria it found decodable emotion information but no verified emotion circuit. The public bundle passed source linting, six software tests, and PDF structural validation. The technical result is deliberately modest: emotion labels were recoverable from internal activations, but the planned causal tests did not verify an emotion circuit. The publication itself is also an experiment in provenance. Releasing the materials can invite useful scrutiny, but visibility does not convert AI output into trustworthy science; only competent audit, replication, and correction can do that.
The report is AI-generated, not peer reviewed, and not independently audited by a domain expert. It tests two checkpoints from one model family, may inherit evaluator and implementation biases, and excludes raw source narratives and generation-level tables from the public bundle. The AI system may have made conceptual, mathematical, statistical, or coding errors. A qualified mechanistic-interpretability researcher should audit the assumptions and implementation, reproduce the results independently, and revise or reject the conclusions as the evidence warrants.