AWS adds Moonshot AI's Kimi K3 model to Amazon Bedrock, with explicit prompt caching support
Moonshot AI's Kimi K3 model became available on Amazon Bedrock (via bedrock-runtime, Invoke/Converse APIs, and OpenAI-compatible Responses/Chat Completions APIs), including a global cross-Region inference profile and a US-only geographic profile. It is the first open-weight model on Bedrock to support explicit prompt caching, letting customers mark reusable prompt prefixes for reduced latency and discounted input-token pricing on cache hits.
Entities: Amazon Web Services, Amazon Bedrock, Moonshot AI, Kimi K3, Kimi K2, OpenCode
1 primary
What happened
AWS added Moonshot AI's Kimi K3 model to Amazon Bedrock, accessible via the bedrock-runtime Invoke/Converse APIs and OpenAI-compatible endpoints, with both a global cross-Region inference profile and a US-only profile. AWS states this is the first open-weight model on Bedrock to support explicit prompt caching, where customers mark reusable prompt prefixes to cut latency and get a discount (roughly 10% on the global profile) on cache-hit input tokens. This is a distribution and availability event: the model itself was already released by Moonshot AI; what changed is that Bedrock customers can now call it directly.
Why it matters
Developers and enterprises already using Bedrock get an easier way to run a large open-weight model for long-context and coding workloads without moving data outside their existing AWS security boundary, which is a real, near-term procurement decision for teams evaluating model options. The prompt-caching feature is concretely actionable now: it directly lowers cost and latency for repeated-prefix workloads like agentic coding tools (OpenCode, Hermes Agent are named as consumers). The broader competitive impact is limited: this does not change the standing of Moonshot AI versus DeepSeek, Qwen, or Mistral in the open-weight race, it just adds one more distribution channel.
What is noise
Moonshot AI's claims of 2.8 trillion parameters and "~2.5x scaling efficiency" over Kimi K2 come from the vendor and are passed through unverified by AWS and by this extraction. AWS's framing of "sustained investment" in open-weight choice is standard vendor packaging: Bedrock has added dozens of open-weight models before, so this is one more addition, not evidence of a strategic shift. The extraction also cites opencode.ai/config.json as primary evidence, which does not match the actual AWS blog source, a sourcing defect worth flagging.
Watch next
- 01Independent benchmark results for Kimi K3 (coding, long-context, agentic tasks) from a third party, not Moonshot AI or AWS, to check the 2.8T parameter and 2.5x efficiency claims
- 02Actual customer adoption signals on Bedrock (usage tier announcements, case studies, or pricing changes tied to Kimi K3) over the next 1-3 months, since availability alone does not mean uptake
- 03Whether other major clouds (Azure, Google Cloud, Oracle) or inference platforms add Kimi K3 with prompt caching, which would show this is a genuine model-level advance rather than an AWS distribution exclusive
Evidence
1 linkedCoverage
1 storyMore capability signals
Full feed →- Deepseek releases V4.1-Flash, an open-source model that sharply cuts KV cache memory and input-processing compute for AI agents10 Sept 202682
- AI systems outperform expert humans in persuasive communication22 Jun 202681
- Anthropic threat report: Claude abused for malware, drone/missile software, mass surveillance, and industrial-scale distillation by Chinese AI labs11 Sept 202680
- WIRED investigation: Flock Safety's AI person-search tools let police run broad description-based surveillance, with weak guardrails against misuse3 Sept 202680