Moonshot AI 的 Kimi K3 登陆 DigitalOcean 推理服务
DigitalOcean 已将 Moonshot AI 的 Kimi K3 接入 Inference Engine,开发者可通过托管式 Serverless Inference 使用该模型,无需自行管理推理基础设施。
Moonshot AI 的 Kimi K3 现已在 DigitalOcean Inference Engine 上线。开发者可以通过 Serverless Inference 按用量调用,也可以借助 Inference Router 将其纳入现有模型组合,并根据成本、延迟或任务类型进行请求路由。
相关介绍将 Kimi K3 定位为面向长上下文和智能体任务的模型,支持最高 100 万 token 的上下文以及原生视觉能力。DigitalOcean 还表示,该模型可用于其专用推理、Agent Development Kit 和智能体相关服务。
所提供材料显示,Serverless Inference 的参考价格为每百万 token 输入 3 美元、输出 15 美元;实际价格、区域和可用性可能发生变化,部署前应以 DigitalOcean 最新文档为准。
来源证据
Moonshot AI's Kimi K3 Launches on DigitalOcean Serverless Inference · Diggdigg.com · supporting# Moonshot AI's Kimi K3 Launches on DigitalOcean Serverless Inference ## Combined views 90.9K 2 posts, first seen 5d ago # Moonshot AI's Kimi K3 Launches on DigitalOcean Serverless Inference ## Reactions from ranked influencers Kimi K3 is now available on@digitalocean's Serverless Inference! Developers can start building with our most capable model in minutes. .@Kimi\_Moonshot K3 from Moonshot AI is now live on DigitalOcean Inference Engine. 1M-token context, native vision, built to run agentic tasks for hours. Supported on Inference Router and model synthesis to maximize intelligence per dollar. No setup, no model ops. Kimi K3 is now available on @digitalocean 's Serverless Inference! Developers can start building with our most capable model in minutes. [...] ## Reactions from the X community @Kimi\_Moonshot @digitalocean Great to see Kimi K3 becoming easier to access. Lowering the barrier for developers usually leads to more experimentation and better applications. @Kimi\_M
Supported Models on DigitalOcean Inferencedocs.digitalocean.com · supportingCopy page as Markdown View page as Markdown DigitalOcean Inference supports more than 70 foundation, embeddings, and reranking models, including OpenAI (GPT-5.x and open-weight gpt-oss), Anthropic Claude, Meta Llama, DeepSeek, Alibaba Qwen, Moonshot AI Kimi, NVIDIA Nemotron, and Z.ai GLM. All text models are served through OpenAI-compatible endpoints, so you can use an existing OpenAI SDK or client by changing the base URL to ` and authenticating with a DigitalOcean API key. For endpoint details, see Serverless Inference Endpoints. Note For pricing information, see the pricing page. [...] ### DeepSeek Models on DigitalOcean DigitalOcean hosts DeepSeek V4 Pro and V4 Flash (with input context windows of up to 1M tokens), DeepSeek V3.2, DeepSeek V3, and DeepSeek R1 Distill Llama 70B. ### Kimi K2.5 and K2.6 on DigitalOcean DigitalOcean hosts Moonshot AI’s Kimi K2.5 and Kimi K2.6 for serverless and dedicated inference, with prompt caching and support for the Chat Completions and Re
DigitalOcean Release Notes - July 2026 Latest Updates - Releasebotreleasebot.io · supportingDigitalOcean logo DigitalOcean ## 27 July DigitalOcean adds Moonshot AI’s Kimi K3 model to Inference for serverless, dedicated inference, ADK, and agents. ### The following Moonshot AI model is now available on DigitalOcean Inference for serverless inference, dedicated inference, Agent Development Kit, and agents: + Kimi K3 For more information, see the Available Models page. Original source All of your release notes in one feed Join Releasebot and get updates from DigitalOcean and hundreds of other software products. Create account Get updates with: Jul 24, 2026 + Date parsed from source: Jul 24, 2026 + First seen by Releasebot: Jul 25, 2026 DigitalOcean logo DigitalOcean ## Now Available: Claude Opus 5 from Anthropic [...] Releasebot Docs Pricing Browse feeds Latest Sign in Create account # DigitalOcean Release Notes Follow Follow DigitalOcean to add their release notes to your feed! 154 release notes curated from 16 source
Under the Hood: Serving Kimi K3 | DigitalOceandigitalocean.com · supporting## What’s next for Kimi K3 on DigitalOcean Kimi K3 is available today on DigitalOcean Inference Engine. Access it through Serverless Inference for fully managed, usage-based inference with no infrastructure to operate, or add it to your existing model mix with the Inference Router to intelligently route requests based on cost, latency, or workload. We’re excited to see what customers build with K3, and we’ll be sharing more soon about how teams are using it in production. Want to see Kimi K3 in action? We built the same AI agent with and without Kimi K3, then compared the results side by side. See what changed, what didn’t, and where Kimi K3 made the biggest difference. Check it out: We Built the Same App Twice With and Without Kimi K3 [...] Standing up a new model, integrating it into DigitalOcean’s Inference Engine, and showcasing its unique attributes on day 0 takes three things: the right hardware, a tuned serving stack, and rigorous verification against Moonshot’s own benchmarks
What's New on DigitalOcean's Inference Enginedigitalocean.com · supportingYou can now access DigitalOcean Serverless Endpoints directly from build.nvidia.com, letting you experiment with elite open‑weight models like Zhipu AI GLM‑5, Moonshot AI Kimi‑K2.5, and MiniMax‑M2.5, and other frontier models from DeepSeek, Meta, and OpenAI. Start building with NVIDIA NemoClaw and the Agent Toolkit on build.nvidia.com, then deploy seamlessly to DigitaOcean’s fully managed Serverless Inference, eliminating infrastructure headaches and giving you more time to focus on shipping AI-powered products. Product Update Now Available: Open-Weight Models on DigitalOcean Serverless Inference You can now run these powerful open-weight models directly on DigitalOcean Serverless Inference, fully integrated with your existing environment: [...] Models available now: DeepSeek-V3.2 DeepSeek-V4-Pro DeepSeek-V4-Flash Kimi K2.6 Kimi K2.5 GLM-5 GLM-5.1 GLM-5.2 gpt-oss-120b Qwen 3.5 MiMo-V2.5 MiMo-V2.5-Pro MiniMax M2.5 Qwen 3 Coder Read the documentation to learn how to str
What Kimi K3 Costs to Run | DigitalOceandigitalocean.com · supportingIf you choose to go with serverless inference, DigitalOcean offers Kimi K3 on the Serverless Inference API as a pay-per-token option at the $3/$15 rate. If you choose to go the self-hosted GPU route, you can use DigitalOcean’s GPU Droplets, Dedicated Inference, or long-term reserved GPUs. With GPU Droplets, you manage the stack, and with Dedicated Inference, DigitalOcean runs a managed vLLM endpoint that you bring the Kimi K3 weights to. The 4-bit native options DigitalOcean provides are the AMD Instinct MI350X and the NVIDIA B300. Availability can change without notice, so be sure to check what is currently available and on which regions before planning your deployment. [...] Kimi K3’s first-party API is $3.00 input / $0.30 cached input / $15.00 output per 1M tokens, with DigitalOcean live at the same $3/$15. The cheapest native-FP4 single node is ~8x MI350X at ~$4.76/GPU-hour. This is approximately $38/hr and $27,800/month on a reserved GPU Droplet that is running whether or not anyo