-
LiveCuratorBench 2026-08-30 Release: Qwen 3.8 Flash and GLM-5.3 Flash
2026-08-30 · Benchmark
Two new Flash-tier models top their harnesses on LiveCuratorBench: Qwen 3.8 Flash scores 20.67/22 on the Qoder harness, the highest score recorded on the benchmark and the only Qoder model past 19, while GLM-5.3 Flash scores 19.67/22 on the opencode harness, the first model there to hold the input-pollution probe in all three runs. A control retest attributes the GLM gain to the model rather than the harness upgrade, and a rubric-calibration audit corrects two published scores (Grok 4.6 to 18.33, HY3 to 18.00).
-
LiveCuratorBench 2026-08-15 Release: GLM-5.3, Grok 4.6, DeepSeek v4 Pro 0813
2026-08-15 · Benchmark
Three new opencode-harness results on LiveCuratorBench: Grok 4.6 scores 18.33/22 after a rubric-calibration audit lowered its initial 19.33, leaving it 1.33 below predecessor Grok 4.5; DeepSeek v4 Pro's stable 0813 release scores 19.00 with zero run-to-run variance, +1.33 over the preview version tested in July; GLM-5.3 scores 18.67/22, tying GLM-5.2 on the same channel with an identical run profile.
-
LiveCuratorBench on the Qoder Harness: Six More Models
2026-08-10 · Benchmark
Six models reran LiveCuratorBench on the Qoder CLI harness: charted one entry per model at its stronger harness, with a separate pairing view for the four models tested on both, the same model can swing by up to 2.7 points across harnesses. Qwen 3.8 Max leads the Qoder round at 18.67/22; DeepSeek v4 Flash is the best value on either harness.
-
Qwen 3.8 Max Tops the Qoder Harness Round on LiveCuratorBench
2026-08-10 · Benchmark
On LiveCuratorBench's Qoder harness round, Qwen 3.8 Max scores 18.67/22—first in the six-model field, ahead of a three-way tie at 18.00—with its edge concentrated in two rubric items no other model held across all three runs.
-
LiveCuratorBench: First Results from a Live Agentic Meta-Knowledge Benchmark
2026-08-01 · Benchmark
LiveCuratorBench is a closed, agentic benchmark that tests whether a model can explore a long practice document and extract reusable methodology while separating it from one-off execution details. Both the data and the ground truth are hand-labeled and never published, so the benchmark cannot be gamed. Across the first thirteen models, Grok 4.5 took first place; DeepSeek v4 Flash finished close behind at a tenth of the cost.
-
DeepSeek v4 Flash (Official): Best Value on LiveCuratorBench
2026-08-01 · Benchmark
On LiveCuratorBench, the official version of DeepSeek v4 Flash scores 19.33/22—second only to Grok 4.5—at $0.027 per run, about a tenth of Grok's cost: the best performance per dollar in the thirteen-model field.
-
On Knowing Things: A Public Misjudgment
2026-07-29 · Opinion
The stock market taught me that most of what I believed had never been tested, because ordinary life never prices a wrong idea. The 2023 claim that China would never catch up in AI failed for the same reason: a static, single-variable model applied to an adaptive system. Knowing a thing takes three layers of method—On Practice for where knowledge comes from, Munger's mental models for checking it, and On Contradiction for reading where it is going.
-
Some of the Agent Era's Biggest Advances Are Conceptual, Not Technical: From MCP to Skills
2026-07-26 · Opinion
Several of the Agent field's biggest recent advances—MCP, skills, the self-iterative agent—were conceptual syntheses, not technical inventions: someone gave a clear name to what everyone was already doing, and the field reorganized around it. Moving from the specific to the abstract is trained deliberately in economics and management, but rarely in STEM—and it is becoming a core research skill.
-
Hello World
2026-07-22 · Opinion
This blog is officially open. Notes on agent systems, AI4Science, and research workflows will appear here.