‹ Experiments

Mentat vs Wolfram Alpha vs Gemini

Sign in

Mentat vs Wolfram Alpha vs Gemini

150 physics questions across 14 domains, each scored against ground truth. Mentat is Aaron's own physics tool server; here a small local model (Qwen 2.5, 7B parameters, running on a home GPU, free) answers by calling it. Every number below is read straight from the result files committed in sigma-ground; nothing is re-scored or hand-picked.

Mentat + Qwen 2.5 7B (local, free)
90.7%
Gemini 2.5 Pro (no tools)
62.0%
Wolfram Alpha (misses re-run)
37.3%
Wolfram Alpha (single pass)
20.0%

Correct / total: Mentat + Qwen 2.5 7B (local, free) 136/150; Gemini 2.5 Pro (no tools) 93/150; Wolfram Alpha (misses re-run) 56/150; Wolfram Alpha (single pass) 30/150. Wolfram Alpha ran on its free tier; "misses re-run" retried the questions it failed to answer the first time.

By domain

DomainMentat + Qwen 2.5 7B (local, free)Gemini 2.5 Pro (no tools)Wolfram Alpha (misses re-run)Wolfram Alpha (single pass)
astrophysics 92% 50% 58% 42%
atomic molecular 75% 50% 62% 50%
classical mechanics advanced 80% 80% 30% 10%
classical mechanics intro 93% 93% 73% 27%
cosmology 100% 25% 25% 12%
electrodynamics advanced 80% 80% 30% 0%
electromagnetism intro 87% 80% 7% 33%
general relativity 90% 50% 20% 0%
mathematical methods 100% 43% 29% 29%
modern physics 100% 75% 33% 8%
nuclear physics 57% 57% 57% 0%
quantum mechanics 100% 8% 42% 17%
thermodynamics statmech 100% 58% 42% 25%
waves optics 100% 83% 17% 17%

Adversarial questions

False premises, impossible requests and trick puzzles, where the right answer is often "that can't be answered as asked". Mentat: 50.0% (9/18); Wolfram Alpha: 0.0% (0/18).

Source: sigma_ground/mcp/benchmark (sigma_ground_scored.json, gemini_scored.json, wolfram_merged_scored.json, wolfram_scored.json, adversarial_sigma_ground_scored.json, adversarial_wolfram_scored.json).