DeepSeek V4 Flash появился в планах ColaOS, но превосходство над GLM-5.2 не подтверждено
DeepSeek V4 Flash упоминается в каталогах моделей и на страницах сравнений как модель с большим контекстом, высокой скоростью и низкой стоимостью. Данные не подтверждают все детали запуска в ColaOS, а стороннее сравнение противоречит заявлению о превосходстве над GLM-5.2 по интеллекту.
DeepSeek V4 Flash указан на некоторых облачных платформах как модель для рассуждений и генерации текста. В опубликованных описаниях говорится о контекстном окне до 1 млн токенов, а также о преимуществах по цене и времени отклика в отдельных сценариях. При этом показатели зависят от поставщика и методики тестирования.
В исходной публикации утверждается, что модель уже добавлена в Token Plan и систему баллов ColaOS и доступна для бесплатного тестирования. Однако представленные материалы в основном состоят из публикаций в соцсетях, видеороликов и агрегаторов. Официального сообщения ColaOS с подробными условиями тарифа или бесплатного доступа нет, поэтому эти сведения требуют дополнительной проверки.
Заявление о превосходстве по интеллекту также вызывает сомнения. В прямом сравнении Artificial Analysis указано, что GLM-5.2 имеет более высокий Intelligence Index, чем DeepSeek V4 Flash 0731. Это противоречит рекламной формулировке, поэтому на данный момент корректнее говорить о конкурентной цене и задержке ответа, а не о подтверждённом лидерстве DeepSeek по качеству.
Источники
DeepSeek V4 Flash 0731 (Reasoning, Max Effort) vs GLM- ...artificialanalysis.ai · supporting| DeepSeek V4 Flash 0731 (Reasoning, Max Effort) | Z AI logoZ AI GLM-5.2 (max) | --- | | Intelligence Index | | | GLM-5.2 (max) is more intelligent than DeepSeek V4 Flash 0731 (Reasoning, Max Effort) | | Price per 1M Tokens | | | DeepSeek V4 Flash 0731 (Reasoning, Max Effort) is cheaper than GLM-5.2 (max) | | Output Speed | 124 tokens/s | 186 tokens/s | GLM-5.2 (max) is faster than DeepSeek V4 Flash 0731 (Reasoning, Max Effort) | | Time to First Token | | | DeepSeek V4 Flash 0731 (Reasoning, Max Effort) responds faster than GLM-5.2 (max) | | Context Window | 1000k tokens~1,500 A4 pages of size 12 Arial font | 1000k tokens~1,500 A4 pages of size 12 Arial font | Both DeepSeek V4 Flash 0731 (Reasoning, Max Effort) and GLM-5.2 (max) have the same sized context window | [...] Reasoning models are indicated by a lightbulb icon The total number of trainable weights and biases in the model, expressed in billions. These parameters are learned during training and determine the model's a
Token Plan Overviewintl.cloud.tencent.com · supporting| | | | | --- --- | | Model Name | Model ID | Model Capability | Length Limits (tokens) | | GLM-5.2 | glm-5.2 glm-5-2 | Deep reasoning, text generation | Context window: 1M Maximum input: 1M Maximum output: 128k | | Kimi-K2.6 | kimi-k2.6 kimi-k-2-6 | Deep reasoning, text generation, image understanding | Context window: 256k Maximum input: 256k Maximum output: 256k | | DeepSeek-V4-Pro (Vendor Direct) | deepseek-v4-pro-202606 | Deep reasoning, text generation | Context window: 1M Maximum input: 1M Maximum output: 384k | | DeepSeek-V4-Flash (Vendor Direct) | deepseek-v4-flash-202605 | Deep reasoning, text generation | Context window: 1M Maximum input: 1M Maximum output: 384k |
GPT-5.6 Luna vs DeepSeek V4 Flash 0423 - AI Model Comparisonopenrouter.ai · supportingInput $0.0882/ M tokens GPT-5.6 Luna Output $0.60/ M tokens DeepSeek V4 Flash Output $0.1764/ M tokens GPT-5.6 Luna Cached input $0.01/ M tokens DeepSeek V4 Flash Cached input $0.01764/ M tokens GPT-5.6 Luna Weighted Average Input $0.03868/ M tokens DeepSeek V4 Flash Weighted Average Input $0.05102/ M tokens GPT-5.6 Luna Cache Write from$0.125/M tokens≤272K $0.125, >272K $0.25 DeepSeek V4 Flash Cache Write — ## Performance GPT-5.6 Luna Latency (p50) 2.60 s DeepSeek V4 Flash Latency (p50) 0.90 s GPT-5.6 Luna Throughput (p50) 51.0 tok/s DeepSeek V4 Flash Throughput (p50) 75.0 tok/s GPT-5.6 Luna Visualize Performance Ready Output will appear here... DeepSeek V4 Flash Visualize Performance Ready Output will appear here... ## Features [...] Cancel Close 40 [...] Ready Output will appear here... ## Features GPT-5.6 Luna Quantization unknown DeepSeek V4 Flash Quantization fp8 GPT-5.6 Luna Max output tokens 128K DeepSeek V4 Flash Ma
DeepSeek 又出手了。最新推出的V4 Flash 0731,在Artificial Analysis 智能 ...facebook.com · supporting## OMP 網路行銷玩家's Post ### OMP 網路行銷玩家 2d · DeepSeek 又出手了。最新推出的 V4 Flash 0731,在 Artificial Analysis 智能指數(綜合九項測試的評分標準)上拿下 50 分,比今年四月推出的上一代 V4 Flash(40 分)大幅跳升 10 分,甚至超越了自家的旗艦 V4 Pro(44 分)。架構和定價與舊版完全相同,完整權重預計未來數週內開放下載。 這個分數已經追近 GPT-5.6 Luna(51 分)及 GLM-5.2(51 分),與 Gemini 3.6 Flash(50 分)同級,距離現時開源模型之王 Kimi K3(57 分)也只差 7 分。更值得留意的是性價比:即使 OpenAI 在同一天為 GPT-5.6 Luna 大幅減價 80%,DeepSeek V4 Flash 0731 透過官方 API 執行同等智能任務的成本,仍然比 GPT-5.6 Luna 低約六成。關鍵在於 DeepSeek 提供高達 98% 的快取命中折扣,遠高於業界普遍的九成水平,令這個模型穩坐「智能對成本」評比中最具吸引力的區間。 實務能力方面同樣進步明顯。在模擬真實工作任務的 GDPval-AA v2 測試中,Elo 評分由上一代的 1189 大幅提升至 1559;一旦完整權重釋出,將成為僅次於 Kimi K3 的第二強開源模型,超越 GLM-5.2。終端操作測試 Terminal-Bench 2.1 上升 17 個百分點至 79%,銀行業務代理測試 τ³-Bench Banking 上升 8 個百分點至 31%。有趣的是,雖然表現全面提升,但執行整套智能指數測試所用的輸出 Token 反而減少 12%,代表新模型不但更聰明,也更省力。 [...] 另一個亮點是「幻覺」問題有實質改善。AA-Omniscience 指數(衡量模型知識可靠度與胡亂作答傾向的分數)由 -23 進步至 -16,但這次進步並非來自答對率提升——準確率其實維持在 37% 不變,而是「亂答」的比例明顯下降:幻覺率由 96% 降至 84%,已經接近 GPT-5.6 Terra(85%)及 Mistral Medium 3.5(82%)的水平。換言之,模型變得更懂得在不確定時選擇不作答,而非硬答錯。 規格方面,1M token
DeepSeek V4 Flash Is Beating Models That Cost 50x MORE!youtube.com · supportingbut remember, it's a little bit more expensive than version 4 Flash. And we see here is GLM 5.2. Deepseek's own version 4 Pro, which is also getting an upgrade. So, keep that in mind. This is probably just a midterm release. This model is probably going to be the one that's going to, you know, change things up a lot. So, I'm excited for this release when it comes out. And then Miniax M3 is in this quadrant as well. But, yeah, that is crazy. And then we have other more intelligence models like Cloud Opus 5, but obviously it's in the area where it's a little bit more expensive compared to all of the models over here. So yeah, as this person says, almost impossible to find one that was both more powerful and cheaper than the Deepseek version for Flash release. Here's the model against 5.6 [...] the same model architecture as the preview version and the official release of Deepseek version 4 Pro not the preview is coming soon. So there's one more model coming from the Deepseek team. So kee
I Re-Tested GPT-5.6-Luna and Deepseek-v4-Flash (New ...youtube.com · supportingare kind of more accurate but still if you see like 775 versus 82.5 that may be the statistical difference. The model is roughly on the same level. So yeah in that table Luna was already pretty high except for Luna low level which was okay. And this is kind of another group of the models which is okay. Starting with Sonnet and then GLM and Gemini Flash and Luna Low Level is here. But look at the price, 10 cents per prompt. And the time is also the best out of those models. Time per prompt. And Deepseek Flash was almost at the very end very cheap, 1 cent per prompt, but it was pretty bad at taking the edge case scenarios into account of the result. In fact, the cheapest model that surprised me recently was 10 cent high3, which also had the cost of 1 cent per prompt via open code, but [...] and it's a win-win both for my audience and for myself and yourself, then let's talk. This is my email and tell me about your product. That's it for this time and see you guys in other [...] 59 commen