Average benchmark score vs. GPT-5.5 (xhigh)
ExoMind: Democratizing Scientific Intelligence via Extended-Mind-Inspired Agentic System
Surpassing Frontier Proprietary Models in Scientific Reasoning and Research with Less Data, Small Model, and Low-cost Training
ExoMind is the first extended-mind-inspired agentic system, integrating systematic data engineering, a scientific interaction framework, and a systematic training strategy to advance scientific reasoning and research.
Frontier-level scientific intelligence at 35B scale
ExoMind consistently improves over the base model on all eight datasets and surpasses leading open and closed-source models on most of them, achieving frontier-level scientific intelligence at 35B scale with roughly 277× fewer parameters than GPT-5.5.
Model-size efficiency vs. GPT-5.5 (xhigh)
GPU-count efficiency vs. GPT-5.5-scale training
Training-data efficiency vs. GPT-5.5-scale training
Intelligence beyond a single LLM
Scientific intelligence demands fine-grained, long-horizon reasoning that is difficult to compress into a single model. Inspired by the Extended Mind Thesis, ExoMind organizes intelligence as an agentic system of an LLM, typed interaction objects, and autonomous interaction processes.
A broader paradigm for scaling intelligence
Deep specialization
Shift specialization from repeatedly modifying model parameters toward designing interaction objects that can be tailored to the operations of a task or domain.
New scaling dimensions
Expand capability through richer interaction objects and deeper autonomous interaction, beyond model size, data volume, and training time.
Stronger representational capacity
Organize the model, external objects, and interaction processes as one system for diverse, fine-grained, and long-horizon tasks.
Building ExoMind
ExoMind is built through three connected systems rather than a single training recipe: data engineering determines what is worth learning, interaction design determines how scientific work is performed, and progressive training aligns the model with that process.
ExoMind starts with a large and diverse scientific problem-answer pool, uses strong models to predict each problem's difficulty and potential tool benefit before trajectory generation, and progressively filters and routes problems into pure-reasoning or interaction-reasoning paths. This yields a compact set of high-quality data with stable learning signals.
ExoMind abstracts key operations in scientific reasoning and research into typed interaction objects under an action-observation contract, then distills verified Chain-of-Interaction trajectories for training. The model learns to select objects, absorb external observations, and iteratively update its reasoning state, allowing reusable interaction patterns to emerge from a finite object set.
ExoMind transforms Qwen3.5-35B-A3B into an agentic reasoner through training-inference consistency and hybrid progressive CoI training. A unified interaction interface limits capability degradation, while staged reasoning and interaction trajectories build intrinsic reasoning, basic interaction, and higher-quality interaction patterns. The full specialization runs on 8 GPUs for one to two days.
Scientific and general intelligence evaluation
Evaluate ExoMind along two complementary dimensions: scientific intelligence against frontier open and proprietary models across eight benchmarks, and general intelligence against the 35B base model across six benchmarks.
Benchmark explorer
What the results show
Scientific research
ExoMind reaches 70.0 on FrontierScience-Research and 84.0 on CMT-Benchmark, substantially exceeding the strongest external results, while remaining competitive on the highly challenging HLE and CritPt benchmarks.
Scientific reasoning
It leads the compared models on AMO-Bench, IMO-AnswerBench, HiPhO, and FrontierScience-Olympiad with 78.0, 92.8, 49.7, and 89.0.
General intelligence
ExoMind improves over the 35B base model on all six general benchmarks, including MMLU-Pro, GPQA-Diamond, MedMCQA, LiveCodeBench, xbench, and τ2-Bench, where the score rises from 32.5 to 74.8.
Evaluation scope
The evaluation covers frontier proprietary and open models, eight scientific reasoning and research benchmarks, and six complementary general intelligence benchmarks.
32 comparison models across nine providers and a broad range of scales.
Open-source model| Provider | Model | Release | Parameters | Architecture | Availability |
|---|---|---|---|---|---|
| OpenAI | GPT-5.5 (xhigh) | 2026-04 | ~9.7T | - | Proprietary |
| GPT-5.4 (xhigh) | 2026-03 | ~2.2T | - | Proprietary | |
| Gemini-3.1-Pro-Preview | 2026-02 | ~9.0T | - | Proprietary | |
| Gemini-3.5-Flash-Thinking | 2026-05 | ~6.6T | - | Proprietary | |
| Gemma-4-31B | 2026-04 | 31B | Dense | Open source | |
| Anthropic | Claude-Opus-4.8 | 2026-05 | ~572B | - | Proprietary |
| Claude-Opus-4.8-Thinking | 2026-05 | ~936B | - | Proprietary | |
| DeepSeek | DeepSeek-V4-Pro (Max) | 2026-04 | 1.6T / A49B | MoE | Open source |
| DeepSeek-V4-Flash (Max) | 2026-04 | 284B / A13B | MoE | Open source | |
| Moonshot AI | Kimi-K3 | 2026-07 | 2.8T / - | MoE | Open source |
| Kimi-K2.7-Code | 2026-06 | 1.1T / A32B | MoE | Open source | |
| Kimi-K2.6 | 2026-04 | 1T / A32B | MoE | Open source | |
| Kimi-K2.5 | 2026-01 | 1T / A32B | MoE | Open source | |
| Zhipu AI | GLM-5.2 | 2026-06 | 744B / A40B | MoE | Open source |
| GLM-5.1 | 2026-04 | 744B / A40B | MoE | Open source | |
| GLM-5 | 2026-02 | 744B / A40B | MoE | Open source | |
| MiniMax | MiniMax-M3 | 2026-06 | 428B / A23B | MoE | Open source |
| MiniMax-M2.7 | 2026-05 | 229B / A10B | MoE | Open source | |
| Qwen | Qwen3.7-Max | 2026-05 | ~685B | - | Proprietary |
| Qwen3.6-Max-Preview | 2026-04 | ~685B | - | Proprietary | |
| Qwen3.6-Plus | 2026-04 | ~524B | - | Proprietary | |
| Qwen3.6-35B-A3B | 2026-04 | 35B / A3B | MoE | Open source | |
| Qwen3.5-397B-A17B | 2026-02 | 397B / A17B | MoE | Open source | |
| Qwen3.5-122B-A10B | 2026-02 | 122B / A10B | MoE | Open source | |
| Qwen3.5-35B-A3B | 2026-02 | 35B / A3B | MoE | Open source | |
| Shanghai AI Lab | Agents-A1 | 2026-06 | 35B / A3B | MoE | Open source |
| P1-VL-235B-A22B | 2026-02 | 235B / A22B | MoE | Open source | |
| P1-VL-30B-A3B | 2026-02 | 30B / A3B | MoE | Open source | |
| P1-235B-A22B | 2025-11 | 235B / A22B | MoE | Open source | |
| P1-30B-A3B | 2025-11 | 30B / A3B | MoE | Open source | |
| SU-01 | 2026-05 | 30B / A3B | MoE | Open source | |
| Intern-S2-Preview | 2026-05 | 35B / A3B | MoE | Open source |
~ indicates an IKP-estimated model size; A denotes activated parameters for MoE models.
Incompressible knowledge probe evaluation
IKP evaluation shows that ExoMind substantially improves the IKP accuracy of the base 35B model and reaches a level close to models with approximately one trillion parameters, suggesting a new dimension for scaling scientific intelligence through interaction.
Emergent scientific interaction patterns
More available interaction objects do not automatically produce stronger performance. The key is whether the model can organize finite atomic capabilities into an observation-guided process: identifying its current uncertainty, selecting an appropriate interaction object, absorbing the returned observation, and updating subsequent reasoning. Through hybrid progressive CoI training, ExoMind composes these objects into reusable verification-seeking, evidence-seeking, and hybrid closed-loop patterns.
FrontierScience-Olympiad #04
Resolve a sign ambiguity through executable verification
Derive the electric field outside a grounded conducting cylinder in a uniform transverse field and verify that the result satisfies the physical boundaries.
- Uncertainty
- A sign-sensitive derivative
The cylindrical-gradient factor makes the sign of the tangential field fragile.
- Interaction
- Code Executor
Differentiate the potential symbolically instead of trusting manual recall.
- Observation
- Both components are returned
The executable result exposes the exact radial and tangential signs.
- Reasoning update
- Verify physical limits
Confirm the surface boundary and recover the applied field at infinity.
Trajectory evidence
Isolate the fragile step
The model starts from a grounded-cylinder potential that satisfies the boundary and far-field conditions, then identifies the tangential derivative as the step most likely to contain a sign error.
Phi = -E0*(r - a**2/r)*cos(theta)
Etheta = -(1/r) * dPhi/dtheta
Differentiate in cylindrical coordinates
Phi = -E0*(r - a**2/r)*sp.cos(theta)
Er = -sp.diff(Phi, r)
Etheta = -(1/r)*sp.diff(Phi, theta)
sp.simplify(Er), sp.simplify(Etheta)
Executable feedback fixes the sign
Er = E0*(a**2 + r**2)*cos(theta)/r**2
Etheta = E0*(a**2 - r**2)*sin(theta)/r**2
Absorb the result into the derivation
Rewriting the returned tangential component confirms the negative sign.
ExoMind then checks that the tangential field vanishes at r = a
and that the field approaches E0 x-hat as r tends to infinity.
CMT-Benchmark #05
Ground a symmetry conclusion in primary literature
Determine which time-reversal and rotational symmetries survive in a twisted bilayer system when the two layers have unequal masses.
- Uncertainty
- Layer inversion remains unresolved
The standard continuum symmetries do not directly settle unequal layer masses.
- Interaction
- Scholar, Search, Browser
Discover the exact-symmetry paper, then move from its abstract to the full source.
- Observation
- Ground T and C3z
The source establishes the relevant continuum-model symmetry reference point.
- Reasoning update
- Apply the mass asymmetry
Preserve T and C3z, while unequal layer masses break the layer-exchanging C2y.
Trajectory evidence
Separate source knowledge from the task-specific condition
ExoMind can partially reason about time reversal and rotational symmetry,
but it is uncertain whether the layer-exchanging C2y operation
survives when the two kinetic terms use different masses.
Search for an exact-symmetry reference
"Bistritzer-MacDonald continuum model symmetries C2y C3z"
"twisted bilayer graphene Hamiltonian time reversal C2y C3z"
Identify the primary source
The results surface Twisted bilayer graphene. III. Interacting Hamiltonian and exact symmetries. Its abstract is relevant but does not contain enough detail to answer the layer-mass question.
Follow the paper trail to arXiv:2009.12376
Browser grounding recovers the paper's continuum-model discussion, which
establishes spinless time reversal and C3z rotational symmetry.
The source supplies a trusted symmetry reference rather than the final answer.
Combine evidence with the unequal-mass Hamiltonian
ExoMind preserves T and C3z from the grounded
reference, then observes that C2y exchanges layers and would
require equal masses. The final decision is therefore T and C3z preserved,
C2y broken.
FrontierScience-Research #19
Close the loop between literature and mathematical verification
Construct a quantum shutter protocol and generalize the one-photon mechanism to an N-photon setting while preserving the required orthogonality.
- Dual uncertainty
- Recover and generalize
The trajectory must establish both the source construction and its N-photon extension.
- Evidence grounding
- Web Search and Browser
Locate the quantum-shutter paper and extract its pre- and post-selected states.
- Executable verification
- Code Executor
Check the overlap that determines whether transmitted branches survive.
- Synthesis
- Close the scientific loop
Join source-grounded states with the verified induction argument.
Trajectory evidence
Recognize two coupled uncertainties
The model must recover the original quantum-shutter construction from a reliable source and independently verify that its generalized transmitted branches are removed after post-selection.
Locate the original construction
quantum shutter pre- and post-selected close all N slits
Aharonov Vaidman quantum shutter N slits
K shutters prevent passage of K photons
Find the source and its central claim
Search returns How One Shutter Can Close N Slits
(quant-ph/0206074), which describes a pre- and post-selected
shutter that closes an arbitrary number of slits for a photon.
Ground the pre- and post-selected states
The paper supplies the state construction needed for the derivation:
|Psi_1> = (sum_i |i> + sqrt(N-1)|N+1>) / sqrt(2N-1)
|Psi_2> = (sum_i |i> - sqrt(N-1)|N+1>) / sqrt(2N-1)
Test the generalized orthogonality condition
inner = (m - D) / (2*N - m)
inner.subs({m: 2, D: 2}) # 0
Use the zero overlap to complete the induction
The overlap vanishes whenever D = m. The first shutter removes
branches occupying m distinct slits; subsequent shutters handle
the remaining branches in descending order. ExoMind absorbs this executable
result into a source-grounded construction for K shutters and K photons.
Related scientific intelligence projects
Scientific reasoning models, benchmarks, and multi-agent systems from the broader open ecosystem.
Citation
Please cite the ExoMind technical report using the following entry.
@misc{exomind2026,
title = {ExoMind: Democratizing Scientific Intelligence via Extended-Mind-Inspired Agentic System},
author = {Peng Ye and Zhuo Liu and Jingqi Ye and Fangchen Yu and Shengji Tang and Yichen Jiang and Haonan He and Zongsheng Cao and Tao Chen and Bo Zhang and Wanli Ouyang and Bowen Zhou and Lei Bai},
year = {2026},
note = {Technical report},
url = {https://github.com/AI4SGI/ExoMind/blob/main/Paper.pdf}
}