Kimi K3 стал доступен на Modal с поддержкой кастомного DFlash
Модель Kimi K3 от Moonshot AI запущена на платформе Modal. Modal добавила кастомную модель спекулятивного декодирования DFlash, настроенную под архитектуру K3, и заявляет о скорости до 460 токенов в секунду.
Открытая модель Kimi K3 от Moonshot AI доступна в Modal с первого дня релиза. Разработчики могут обращаться к ней через Shared API с оплатой по числу токенов или использовать Auto Endpoint для выделенных вычислительных ресурсов. По данным Modal, интеграция выполнена совместно с Moonshot AI и vLLM.
Для ускорения генерации используется кастомная модель DFlash, обученная с учетом архитектуры K3. Спекулятивное декодирование предварительно предлагает возможные продолжения, после чего целевая модель их проверяет, что может сократить задержку при генерации больших объемов текста.
Согласно опубликованным данным, Kimi K3 имеет около 2,8 трлн параметров, встроенную поддержку изображений и контекстное окно на 1 млн токенов. Заявление Modal о повышении скорости без потери качества является заявлением поставщика и требует проверки по независимым тестам.
Источники
Kimi K3 by Moonshot now available on Modaldaily.dev · supportingP ro motedb y E t h i c alAd s MongoDB Atlas is the vector database developers prefer. Get started building for free. Advertise here Image 6: PixelImage 7: Pixel #### Would you recommend this post? Copy link WhatsApp Facebook X New Squad Copy link Share with your friends #### Table of contents Open, frontier, fast. Pick three.Drafting for delta attentionRun Kimi K3 on Modal #### Best discussions ##### Linux 7.2 finally drops support for a 44-year-old graphics card 12 Comments ##### New IronWorm malware hits 36 packages in npm supply-chain attack 21 Comments ##### Highlights from Git 2.55 11 Comments I'm feeling lucky © 2026 Daily Dev Ltd. Guidelines Explore Tags Sources Squads Leaderboard Image 8 🇺🇸 [...] Image 8 🇺🇸 ## daily.dev is the fastest growing developer platform in The United States! ### We know how hard it is to be a developer. It doesn't have to be. Personalized news feed, dev community and search, much better tha
Kimi K3 Model Overview: 2.8T Parameters, MXFP4 ...huggingface.co · supporting## Model Card Summary | Field | Value | --- | | Developer | Moonshot AI | | Model type | Autoregressive Mixture-of-Experts (MoE) transformer with native vision | | Total parameters | 2.8 trillion | | Active parameters | ~50B equivalent (16/896 experts per token) | | Context length | 1,000,000 tokens | | Training precision | Mixed (MXFP4 weights, MXFP8 activations) | | Modalities | Text + Vision (native, not adapter-based) | | Reasoning | Always-on thinking mode | | License | Open weights (license TBD with weight release) | | Release date | July 16, 2026 (API); July 27, 2026 (weights) | ## Architecture Kimi K3 introduces several architectural innovations that collectively yield approximately 2.5x scaling efficiency improvement over its predecessor K2. ### Kimi Delta Attention (KDA) [...] Moonshot AI publicly released Kimi K3 on July 16, 2026, with full open-source weights promised by July 27. At 2.8 trillion parameters, it is the first open-source model to reach the 3-trillion-para
Kimi K3 by Moonshot now available on Modalmodal.com · supportingUser avatar Adam Azzam Member of Product Staff Today Moonshot released Kimi K3, a 2.8 trillion parameter multimodal model with a 1M token context window and native vision. And Modal runs it at 460 tokens per second, on release day. We partnered with Moonshot and vLLM on day zero support, making K3 available with token-based pricing on our Shared API, and as an Auto Endpoint for dedicated capacity, alongside a custom-trained DFlash speculator tuned to K3's architecture. Try it now or read on for why we think this model, and its architecture, matter. ## Frontier, open, fast. Pick three. Kimi K3 is the strongest open model on public intelligence indexes, fourth overall in a leaderboard dominated by closed source models. [...] For endpoints on Modal, we took this even further with Day 0 support for a custom DFlash speculator tuned to K3's shape. K3 generates a large number of tokens per task, which means most of the time a user spends waiting is decode time, and decode is what specu
Kimi K3 Is Here: Efficient Day-0 Support on vLLMvllm.ai · supportingInferact has also trained and open-sourced a DSpark speculator for Kimi K3. Enable it by adding the following option to the serve command: ``` --speculative-config '{"model":"Inferact/Kimi-K3-DSpark","method":"dspark","num_speculative_tokens":7,"attention_backend":"FLASHINFER_MLA","draft_sample_method":"probabilistic","rejection_sample_method":"block"}' --speculative-config '{"model":"Inferact/Kimi-K3-DSpark","method":"dspark","num_speculative_tokens":7,"attention_backend":"FLASHINFER_MLA","draft_sample_method":"probabilistic","rejection_sample_method":"block"}' ``` [...] vLLM LogovLLM Logo # Kimi K3 Is Here: Efficient Day-0 Support on vLLM 24 min read vLLM Team and Inferact #models#performance#prefix caching#multimodal We're thrilled to announce efficient day-0 vLLM support for Kimi K3, one of the most powerful open-weight models ever released. Last week, we previewed the production-scale integration work for Kimi K3; today, Moonshot AI's weights are public and the support is l
How to build a day-0 API for Kimi K3baseten.co · supporting## Build with Kimi K3 on Baseten We’re excited to offer day-0 access via Model APIs, and look forward to continuing to optimize our implementation of this model to achieve the highest standards in performance and reliability. The Moonshot AI team’s Kimi K3 announcement includes a number of interesting tests for the model, including coding tasks like kernel optimization and vision-in-the-loop game development, research tasks, and agentic tasks like video editing and knowledge work. On Tuesday, July 28 at 11AM Pacific Time, we are hosting an executive briefing on Kimi K3 use cases with Joey Zwicker, who leads all forward-deployed engineering at Baseten, and Philip Kiely, author of Inference Engineering. [...] Having the model up and running from previous milestones is a dependency for this work. For example, training a speculator model using a method like DSpark, DFlash, or EAGLE-3 requires generating hidden states from the target model (Kimi K3) using a set of prompts that resemble ex
Kimi.ai (@Kimi_Moonshot) / Xx.com · supportingKimi.ai @Kimi\_Moonshot Jul 27 Kimi K3 is now live on @togethercompute! Happy to have Together AI as our day0 launch partner, giving developers immediate access to K3 through high-throughput inference optimized for coding agents and production workloads. A huge thank-you to the Together AI team for user avatar Together AI @togethercompute Jul 27 Kimi K3 is now live on Together AI. We’re proud to be a Day 0 launch partner for @Kimi\_Moonshot’s open frontier model, built for long-running agentic workflows across code, tools, vision, and research. 112K user avatar Kimi.ai @Kimi\_Moonshot Jul 27 Kimi K3 is now available on @digitalocean 's Serverless Inference! Developers can start building with our most capable model in minutes. [...] 00:36 user avatar Jul 27 .@Kimi\_Moonshot K3 from Moonshot AI is now live on DigitalOcean Inference Engine. 1M-token context, native vision, built to run agentic tasks for hours. Supported on Inference