DeepSeek V4.1 Currently 1500x Cheaper Than Usual
DeepSeek V4.1 Flash is currently listed at about 1,500 times cheaper than DeepSeek’s usual off-peak input rate on some third-party hosts. On OpenRouter, Relace shows the model at $0.0001 per million input tokens, while DeepSeek’s own off-peak input price is $0.15 per million.
That gap is what sparked a Reddit thread on r/DeepSeek titled around Relace and Open Inference’s “shockingly low” listed prices. Open Inference’s OpenRouter listing sits near $0.00011 per million input tokens and $0.36 per million output tokens.
The 1,500x figure applies to input pricing only. Relace still lists output at $0.60 per million tokens, which matches DeepSeek’s own off-peak output rate. So the bargain is on the input side of those provider listings, not a full rewrite of every token price.
How The Listed Prices Compare
DeepSeek’s published off-peak schedule for V4.1 Flash is $0.15 per million input tokens and $0.60 per million output tokens, with peak rates doubling both figures. Cache hits drop input cost further on DeepSeek’s own API.
Relace’s OpenRouter row currently shows $0.0001 / $0.60 per million for input and output. Open Inference shows about $0.00011 / $0.36. Against DeepSeek’s $0.15 off-peak input rate, Relace’s $0.0001 input list price is roughly 1,500 times lower.
Those figures are provider list prices on aggregators such as OpenRouter. Listed rates can change, and real bills also depend on cache hits, routing, latency, and how many output tokens a job uses.
What DeepSeek V4.1 Flash Is
DeepSeek V4.1 Flash is a sparse mixture-of-experts model released around September 10, 2026. It is the first model built on DeepSeek’s Causal Encoder-Decoder architecture. The company describes a 552 billion parameter backbone that activates about 8 billion parameters on input and 16 billion on output.
The model is aimed at coding, terminal work, computer-use agents, and long-context jobs. OpenRouter lists a 1 million token context window. DeepSeek also says compressed KV caching cuts cache memory versus the previous Flash generation, which matters for agent workflows that reuse long prompts.
ProPakistani earlier covered the DeepSeek V4.1 Flash launch, including its lower official pricing tier and native vision support. The new attention is less about another official cut and more about how cheap some third-party endpoints are advertising the same model.
Why Developers Are Watching The Listings
For high-volume input workloads, a move from $0.15 to $0.0001 per million tokens would be dramatic if the endpoint stays stable and the output quality holds up. Output cost still sits near DeepSeek’s off-peak $0.60 on Relace, so chatty or long-answer jobs will not see the same 1,500x savings.
Open Inference’s lower output list price ($0.36) is a separate pull, though recent OpenRouter snapshots have also shown weaker uptime and slower throughput on that route compared with Relace. Developers comparing hosts still need to weigh price against reliability and speed.
DeepSeek has also been covered for earlier pricing moves, including when the company raised prices by up to 4x. The Relace and Open Inference listings sit at the opposite extreme: third-party rows that undercut DeepSeek’s own input sticker by orders of magnitude.
For now, the practical takeaway is narrow and checkable. On current OpenRouter listings, Relace and Open Inference advertise DeepSeek V4.1 Flash input rates near $0.0001 per million tokens, about 1,500 times below DeepSeek’s usual $0.15 off-peak input price, while output pricing stays much closer to normal.
Source: r/DeepSeek on Reddit; OpenRouter
See more ProPakistani stories in Google Search and Top Stories.
