Mentat vs Wolfram Alpha vs Gemini
Sign inMentat vs Wolfram Alpha vs Gemini
150 physics questions across 14 domains, each scored against ground truth. Mentat is Aaron's own physics tool server; here a small local model (Qwen 2.5, 7B parameters, running on a home GPU, free) answers by calling it. Every number below is read straight from the result files committed in sigma-ground; nothing is re-scored or hand-picked.
Correct / total: Mentat + Qwen 2.5 7B (local, free) 136/150; Gemini 2.5 Pro (no tools) 93/150; Wolfram Alpha (misses re-run) 56/150; Wolfram Alpha (single pass) 30/150. Wolfram Alpha ran on its free tier; "misses re-run" retried the questions it failed to answer the first time.
By domain
| Domain | Mentat + Qwen 2.5 7B (local, free) | Gemini 2.5 Pro (no tools) | Wolfram Alpha (misses re-run) | Wolfram Alpha (single pass) |
|---|---|---|---|---|
| astrophysics | 92% | 50% | 58% | 42% |
| atomic molecular | 75% | 50% | 62% | 50% |
| classical mechanics advanced | 80% | 80% | 30% | 10% |
| classical mechanics intro | 93% | 93% | 73% | 27% |
| cosmology | 100% | 25% | 25% | 12% |
| electrodynamics advanced | 80% | 80% | 30% | 0% |
| electromagnetism intro | 87% | 80% | 7% | 33% |
| general relativity | 90% | 50% | 20% | 0% |
| mathematical methods | 100% | 43% | 29% | 29% |
| modern physics | 100% | 75% | 33% | 8% |
| nuclear physics | 57% | 57% | 57% | 0% |
| quantum mechanics | 100% | 8% | 42% | 17% |
| thermodynamics statmech | 100% | 58% | 42% | 25% |
| waves optics | 100% | 83% | 17% | 17% |
Adversarial questions
False premises, impossible requests and trick puzzles, where the right answer is often "that can't be answered as asked". Mentat: 50.0% (9/18); Wolfram Alpha: 0.0% (0/18).
Source: sigma_ground/mcp/benchmark (sigma_ground_scored.json, gemini_scored.json, wolfram_merged_scored.json, wolfram_scored.json, adversarial_sigma_ground_scored.json, adversarial_wolfram_scored.json).