Control
0.2466
Single-turn baseline
CONEXUS reports controlled experiments and internal computational benchmarks with their conditions, baselines, and limitations. The evidence supports specific findings. It does not justify universal claims about every model, algorithm, or scientific domain.
The study tested Gemini 3.1 Pro Preview on an Alternative Uses Task, with 50 independent runs per condition, temperature 0.7, 16,000 maximum output tokens, and local BGE embeddings for the semantic-distance measurement.
Control
0.2466
Single-turn baseline
Neutral
0.2219
Analytical multi-turn prompt
Token-only
0.2258
Emoji exposure without the architecture
CONEXUS
0.2929
Complete contradiction-holding sequence
d = 3.7824
Large run-level standardized mean difference in the tested configuration.
2.97e-32
The bootstrap interval for the mean difference was [+0.063467, +0.078094].
p = 0.3612
No statistically detectable difference from the neutral condition was found in this comparison.
Precision strengthens the finding. It does not diminish it.
The full CONEXUS condition produced the highest run-level mean semantic distance in this experiment.
The difference between the neutral and CONEXUS conditions was large in the tested configuration: Cohen's d = 3.7824.
The neutral-to-CONEXUS Welch test returned p = 2.97e-32, and the bootstrap confidence interval for the mean difference excluded zero.
The token-only condition was not statistically distinguishable from the neutral condition: p = 0.3612, with a small effect estimate.
The longer neutral prompt compressed rather than expanded the measured search behavior, weighing against prompt length as the explanation.
One model family and one divergent-thinking task were used in the reported four-arm study.
Semantic distance is a behavioral measurement, not a general measure of intelligence, truth, creativity, or consciousness.
The 39.9242% idea-level variance difference is descriptive; its Levene variance test was not significant at p = 0.304333.
The causal result supports the tested prompt architecture under these conditions. Broader generalization requires additional models, tasks, preregistration, and independent replication.
The locked optimization sweep contains 30,800 controlled trials. Additional domain studies test the same strategic-elimination idea in different search spaces. Each result belongs to its own objective, baseline, and configuration.
Important: a 561% relative success-rate difference in protein folding is not the same quantity as an 89.3% routing improvement or a 27.8% gate reduction. These numbers should be read within their own experiments, not combined into one universal score.
2,000 trials
Approximately 80% relative improvement in the stated comparison
Internal benchmark against the documented Monte Carlo baseline
4,000 trials
25.8% success versus 3.9%, approximately 561% relative improvement
Largest reported relative gap in this research portfolio
Scale series trials
Larger relative gaps were reported at larger tested instances
Benchmark-specific trend, not a universal scaling law
250 trials
Up to 89.3% improvement at the largest tested scale
Compared with the stated routing baseline and configuration
300 trials
Reported accuracy gains ranged from 3.8% to 8.4%
Internal search benchmark; external replication remains needed
5,000 trials
27.8% gate reduction and 3.7% fidelity gain were reported
Simulator-based comparison under the documented compilation setup
In several CONEXUS benchmark series, the relative advantage over the chosen baseline increased at larger tested scales. That is the phenomenon CONEXUS calls complexity inversion.
Establishing a general scaling law would require preregistered experiments, stronger competing methods, multiple independent implementations, and replication outside the CONEXUS team.
Observed
Larger relative gaps in selected benchmark series as tested scale increased.
Not yet established
A universal rule that the Forgetting Engine improves with every form of complexity or defeats all conventional algorithms.
An exploratory analysis retained three anomalous signals from public catalog data for further review. They are not presented as independently confirmed exoplanet discoveries.
Retained by the exploratory anomaly-ranking process for follow-up analysis. Candidate status does not establish a planetary interpretation.
Retained by the exploratory anomaly-ranking process for follow-up analysis. Candidate status does not establish a planetary interpretation.
Retained by the exploratory anomaly-ranking process for follow-up analysis. Candidate status does not establish a planetary interpretation.
The strategic-retention approach can surface and preserve anomalous candidates that might otherwise be eliminated early in a ranking pipeline.
It does not independently validate the candidates as planets, establish a false-positive rate for discovery, or substitute for domain-expert astronomical confirmation.
The public materials provide methods, reported results, and source paths for technical inspection. Availability of a report is not a substitute for independent replication or peer review.
CONEXUS will continue separating demonstrated results from research hypotheses, product concepts, and future applications.
Request Technical MaterialsVisual Evidence Library
These supporting visuals summarize the controlled comparisons behind CONEXUS calibration research. Select either image to open it at full size.

A technical comparison of control, token priming, neutral logical prompting, and the full CONEXUS paradox architecture.

A visual explanation of the controlled evidence separating paradox-holding architecture from prompt length or token exposure.