Read the blog.

Engineering deep-dives, product updates, and the occasional opinion.

Placeholder cover for the DeepSeek V4.1 Flash versus GPT-6 Astra token-efficiency benchmark.
LatestEngineering

Benchmarking cost per correct answer: DeepSeek V4.1 Flash vs GPT-6 Astra

Both models solved one-shot agent and coding tasks. DeepSeek cost $0.00048 per solve versus Astra's $0.017, even though it generated 4.4 times as many output tokens. The suite tests retrieval, coding, structured output, and instruction following, not general intelligence. Will DeepSeek V4.1 Flash replace USA AI Lab models for AI agent use?

September 202616 min readRead more