DeepSeek V4-Flash Now Available on Qiniu Cloud: Claim 3 Million Free Tokens
On July 31, 2026, DeepSeek officially released the V4-Flash version. According to public benchmark data, this version significantly outperforms its own DeepSeek-V4-Pro preview and surpasses GLM-5.2. However, benchmark scores are for reference only; actual performance still needs verification through multi-scenario testing.

The price of DeepSeek-V4-Flash is indeed very affordable, and the speed is fast. Currently, in addition to the official DeepSeek channels, Qiniu Cloud has also become the first to support the official DeepSeek-V4-Flash version. New users receive 3 million tokens upon registration, which can be claimed if needed.
Claim 3 Million Tokens from Qiniu Cloud
Qiniu Cloud 3 Million Token Claim Address: https://s.qiniu.com/ZjmEji
Or scan the code to claim:

Using DeepSeek-V4-Flash Large Model on Qiniu Cloud
After registering and authenticating, create and copy the KEY in the Qiniu Cloud backend: https://portal.qiniu.com/ai-inference/api-key

Then customize the API endpoint in your tool. It is compatible with both OpenAI and Anthropic formats, with the corresponding endpoint addresses as follows:
- OpenAI BaseURL:
https://api.qnaigc.com/v1, full address:https://api.qnaigc.com/v1/chat/completions - Anthropic BaseURL:
https://api.qnaigc.com, full address:https://api.qnaigc.com/v1/messages
Note: The complete model ID for the DeepSeek V4 Flash official version is
deepseek/deepseek-v4-flash-20260731.
Xiaoz integrated the deepseek/deepseek-v4-flash-20260731 official version into Cherry Studio for testing. The speed is truly impressive, reaching 170 tokens/s.

Additionally, besides providing DeepSeek models, Qiniu Cloud also offers other models such as GLM, Kimi, and Minimax for invocation.

Conclusion
The DeepSeek V4-Flash official version shows impressive benchmark scores, with actual test speeds reaching 170 tokens/s and an affordable pricing structure. Qiniu Cloud was the first to integrate it and provides 3 million free tokens, compatible with mainstream API formats, significantly lowering the threshold for developers to try it out. Although actual performance still requires validation across multiple scenarios, its high cost-performance ratio and convenient access provide a new option for AI application deployment, worth trying.