DeepSeek V4-Flash Now Available on Qiniu Cloud: Claim 3 Million Free Tokens

DeepSeek V4-FlashQiniu Cloud AI3 million free tokensAI inference APICherry Studio integration
Published·Modified·

On July 31, 2026, DeepSeek officially released the V4-Flash version. According to public benchmark data, this version significantly outperforms its own DeepSeek-V4-Pro preview and surpasses GLM-5.2. However, benchmark scores are for reference only; actual performance still needs verification through multi-scenario testing.

ds4d8bbe87a7.jpg

The price of DeepSeek-V4-Flash is indeed very affordable, and the speed is fast. Currently, in addition to the official DeepSeek channels, Qiniu Cloud has also become the first to support the official DeepSeek-V4-Flash version. New users receive 3 million tokens upon registration, which can be claimed if needed.

Claim 3 Million Tokens from Qiniu Cloud

Qiniu Cloud 3 Million Token Claim Address: https://s.qiniu.com/ZjmEji

Or scan the code to claim:

七牛 AI 推理邀请海报.png

Using DeepSeek-V4-Flash Large Model on Qiniu Cloud

After registering and authenticating, create and copy the KEY in the Qiniu Cloud backend: https://portal.qiniu.com/ai-inference/api-key

CleanShot 2026-08-01 at 11.50.59@2x.png

Then customize the API endpoint in your tool. It is compatible with both OpenAI and Anthropic formats, with the corresponding endpoint addresses as follows:

  • OpenAI BaseURL: https://api.qnaigc.com/v1, full address: https://api.qnaigc.com/v1/chat/completions
  • Anthropic BaseURL: https://api.qnaigc.com, full address: https://api.qnaigc.com/v1/messages

Note: The complete model ID for the DeepSeek V4 Flash official version is deepseek/deepseek-v4-flash-20260731.

Xiaoz integrated the deepseek/deepseek-v4-flash-20260731 official version into Cherry Studio for testing. The speed is truly impressive, reaching 170 tokens/s.

CleanShot 2026-08-01 at 11.55.00@2x.png

Additionally, besides providing DeepSeek models, Qiniu Cloud also offers other models such as GLM, Kimi, and Minimax for invocation.

CleanShot 2026-08-01 at 11.59.08@2x.png

Conclusion

The DeepSeek V4-Flash official version shows impressive benchmark scores, with actual test speeds reaching 170 tokens/s and an affordable pricing structure. Qiniu Cloud was the first to integrate it and provides 3 million free tokens, compatible with mainstream API formats, significantly lowering the threshold for developers to try it out. Although actual performance still requires validation across multiple scenarios, its high cost-performance ratio and convenient access provide a new option for AI application deployment, worth trying.