公告
数据公告

QQ群和tg群已经启用,欢迎加入。公开信息来源均审核后发布;请结合来源、库存和更新时间判断。

社群与联系Telegram 群点击加入Telegram 频道点击订阅联系我们tgAIPricedb交流群979789483
返回资讯列表
model_api

Kimi K3 在 Modal 上线,配备定制 DFlash 推测解码支持

Moonshot AI 的 Kimi K3 已在 Modal 上线。Modal 表示,其与 Moonshot AI 和 vLLM 合作提供首日支持,并使用针对 K3 架构训练的定制 DFlash 推测模型提升推理速度。

91% VERIFIED

Moonshot AI 的开放权重模型 Kimi K3 已在 Modal 上提供服务,用户可通过 Modal Shared API 使用按 token 计费的接口,也可以选择面向专用容量的 Auto Endpoint。Modal 称,该模型在发布日即可获得支持,实测推理速度最高可达每秒 460 个 token。

为适配 K3 的架构,Modal 与合作方部署了定制 DFlash 推测模型。推测解码通过预测后续 token 来减少目标模型的解码开销,尤其适合 K3 这类单次任务生成量较大的模型。

相关资料显示,Kimi K3 拥有约 2.8 万亿参数、原生视觉能力和 100 万 token 上下文窗口,并已在 Together AI、DigitalOcean 等平台提供服务。Modal 关于“无质量损失”的表述属于其部署方案的性能主张,仍应结合独立基准进行评估。

来源证据

Kimi K3 by Moonshot now available on Modaldaily.dev · supporting

P r​o m​o​te​d​b y​ E t h i c a​l​A​d s​ MongoDB Atlas is the vector database developers prefer. Get started building for free. Advertise here Image 6: PixelImage 7: Pixel #### Would you recommend this post? Copy link WhatsApp Facebook X New Squad Copy link Share with your friends #### Table of contents Open, frontier, fast. Pick three.Drafting for delta attentionRun Kimi K3 on Modal #### Best discussions ##### Linux 7.2 finally drops support for a 44-year-old graphics card 12 Comments ##### New IronWorm malware hits 36 packages in npm supply-chain attack 21 Comments ##### Highlights from Git 2.55 11 Comments I'm feeling lucky © 2026 Daily Dev Ltd. Guidelines Explore Tags Sources Squads Leaderboard Image 8 🇺🇸 [...] Image 8 🇺🇸 ## daily.dev is the fastest growing developer platform in The United States! ### We know how hard it is to be a developer. It doesn't have to be. Personalized news feed, dev community and search, much better tha

Kimi K3 Model Overview: 2.8T Parameters, MXFP4 ...huggingface.co · supporting

## Model Card Summary | Field | Value | --- | | Developer | Moonshot AI | | Model type | Autoregressive Mixture-of-Experts (MoE) transformer with native vision | | Total parameters | 2.8 trillion | | Active parameters | ~50B equivalent (16/896 experts per token) | | Context length | 1,000,000 tokens | | Training precision | Mixed (MXFP4 weights, MXFP8 activations) | | Modalities | Text + Vision (native, not adapter-based) | | Reasoning | Always-on thinking mode | | License | Open weights (license TBD with weight release) | | Release date | July 16, 2026 (API); July 27, 2026 (weights) | ## Architecture Kimi K3 introduces several architectural innovations that collectively yield approximately 2.5x scaling efficiency improvement over its predecessor K2. ### Kimi Delta Attention (KDA) [...] Moonshot AI publicly released Kimi K3 on July 16, 2026, with full open-source weights promised by July 27. At 2.8 trillion parameters, it is the first open-source model to reach the 3-trillion-para

Kimi K3 by Moonshot now available on Modalmodal.com · supporting

User avatar Adam Azzam Member of Product Staff Today Moonshot released Kimi K3, a 2.8 trillion parameter multimodal model with a 1M token context window and native vision. And Modal runs it at 460 tokens per second, on release day. We partnered with Moonshot and vLLM on day zero support, making K3 available with token-based pricing on our Shared API, and as an Auto Endpoint for dedicated capacity, alongside a custom-trained DFlash speculator tuned to K3's architecture. Try it now or read on for why we think this model, and its architecture, matter. ## Frontier, open, fast. Pick three. Kimi K3 is the strongest open model on public intelligence indexes, fourth overall in a leaderboard dominated by closed source models. [...] For endpoints on Modal, we took this even further with Day 0 support for a custom DFlash speculator tuned to K3's shape. K3 generates a large number of tokens per task, which means most of the time a user spends waiting is decode time, and decode is what specu

Kimi K3 Is Here: Efficient Day-0 Support on vLLMvllm.ai · supporting

Inferact has also trained and open-sourced a DSpark speculator for Kimi K3. Enable it by adding the following option to the serve command: ``` --speculative-config '{"model":"Inferact/Kimi-K3-DSpark","method":"dspark","num_speculative_tokens":7,"attention_backend":"FLASHINFER_MLA","draft_sample_method":"probabilistic","rejection_sample_method":"block"}' --speculative-config '{"model":"Inferact/Kimi-K3-DSpark","method":"dspark","num_speculative_tokens":7,"attention_backend":"FLASHINFER_MLA","draft_sample_method":"probabilistic","rejection_sample_method":"block"}' ``` [...] vLLM LogovLLM Logo # Kimi K3 Is Here: Efficient Day-0 Support on vLLM 24 min read vLLM Team and Inferact #models#performance#prefix caching#multimodal We're thrilled to announce efficient day-0 vLLM support for Kimi K3, one of the most powerful open-weight models ever released. Last week, we previewed the production-scale integration work for Kimi K3; today, Moonshot AI's weights are public and the support is l

How to build a day-0 API for Kimi K3baseten.co · supporting

## Build with Kimi K3 on Baseten We’re excited to offer day-0 access via Model APIs, and look forward to continuing to optimize our implementation of this model to achieve the highest standards in performance and reliability. The Moonshot AI team’s Kimi K3 announcement includes a number of interesting tests for the model, including coding tasks like kernel optimization and vision-in-the-loop game development, research tasks, and agentic tasks like video editing and knowledge work. On Tuesday, July 28 at 11AM Pacific Time, we are hosting an executive briefing on Kimi K3 use cases with Joey Zwicker, who leads all forward-deployed engineering at Baseten, and Philip Kiely, author of Inference Engineering. [...] Having the model up and running from previous milestones is a dependency for this work. For example, training a speculator model using a method like DSpark, DFlash, or EAGLE-3 requires generating hidden states from the target model (Kimi K3) using a set of prompts that resemble ex

Kimi.ai (@Kimi_Moonshot) / Xx.com · supporting

Kimi.ai @Kimi\_Moonshot Jul 27 Kimi K3 is now live on @togethercompute! Happy to have Together AI as our day0 launch partner, giving developers immediate access to K3 through high-throughput inference optimized for coding agents and production workloads. A huge thank-you to the Together AI team for user avatar Together AI @togethercompute Jul 27 Kimi K3 is now live on Together AI. We’re proud to be a Day 0 launch partner for @Kimi\_Moonshot’s open frontier model, built for long-running agentic workflows across code, tools, vision, and research. 112K user avatar Kimi.ai @Kimi\_Moonshot Jul 27 Kimi K3 is now available on @digitalocean 's Serverless Inference! Developers can start building with our most capable model in minutes. [...] 00:36 user avatar Jul 27 .@Kimi\_Moonshot K3 from Moonshot AI is now live on DigitalOcean Inference Engine. 1M-token context, native vision, built to run agentic tasks for hours. Supported on Inference