Science · AI-generated study

AI Emotion-Circuits Report

An explicitly AI-generated and unverified mechanistic-interpretability experiment on emotion-related representations in two small language models.

Compendium article 049 Revision 0.6 · July 2026

This record is intentionally unusual: the research producer is an AI system, while the human who initiated and released the work does not claim to understand or independently verify its technical substance. The disclosure is part of the artifact, not a footnote.

The project takes the form of an end-to-end mechanistic-interpretability study generated and executed by OpenAI GPT-5.6 Sol in a high-reasoning Codex session. It investigates linear decodability, activation steering, necessity, controls, and sparse-component localization in Qwen2.5 0.5B and 1.5B Instruct. Its purpose is to test whether emotion-related representations in sub-2B language models behave as causally operative mechanisms while documenting what an autonomous AI research workflow can and cannot establish.

The question behind AI Emotion-Circuits Report

The experiment began with a genuine research question about whether emotion-related representations in a small language model could be identified through a repeatable process. Josiah then asked GPT-5.6 Sol, running in Codex at high reasoning, to carry out the entire study. That created a second question alongside the scientific one: how should a polished result be published when its human custodian cannot responsibly claim authorship or validation? Mechanistic-interpretability researchers and readers interested in autonomous AI research can inspect or reproduce the artifact. Nobody should treat its technical claims as established without independent expert audit and replication.

How AI Emotion-Circuits Report took shape

The AI system turned the question into operational definitions and prespecified tests, assembled datasets, implemented and tested an experimental pipeline, ran representation probes and causal interventions across two Qwen checkpoints, analyzed the outputs, and compiled the report. Before public release, the title page, repository, metadata, and website framing were revised to name the AI system as the research producer and state Josiah's narrower role. Josiah initiated the question, provided the computing environment, observed the workflow, challenged the authorship framing, and authorized release. He does not claim technical authorship, comprehension of the mathematical and statistical details, or independent verification. GPT-5.6 Sol through Codex performed the technical research and writing.

What the evidence supports

The study reports held-out balanced probe accuracies of 0.601 and 0.701, but mean causal steering effects of only 0.0133 and 0.0015. Under its prespecified criteria it found decodable emotion information but no verified emotion circuit. The public bundle passed source linting, six software tests, and PDF structural validation. The technical result is deliberately modest: emotion labels were recoverable from internal activations, but the planned causal tests did not verify an emotion circuit. The publication itself is also an experiment in provenance. Releasing the materials can invite useful scrutiny, but visibility does not convert AI output into trustworthy science; only competent audit, replication, and correction can do that.

The report is AI-generated, not peer reviewed, and not independently audited by a domain expert. It tests two checkpoints from one model family, may inherit evaluator and implementation biases, and excludes raw source narratives and generation-level tables from the public bundle. The AI system may have made conceptual, mathematical, statistical, or coding errors. A qualified mechanistic-interpretability researcher should audit the assumptions and implementation, reproduce the results independently, and revise or reject the conclusions as the evidence warrants.

Published artifact

Emotion Representations Without Verified Emotion Circuits

An 11-page AI-generated experimental report with explicit provenance, methods, aggregate results, limitations, and a warning that no independent expert has verified the work.