Signum
Feed
Useful signal18 Sept 2026high confidence

AWS adds Moonshot AI's Kimi K3 model to Amazon Bedrock, with explicit prompt caching support

Moonshot AI's Kimi K3 model became available on Amazon Bedrock (via bedrock-runtime, Invoke/Converse APIs, and OpenAI-compatible Responses/Chat Completions APIs), including a global cross-Region inference profile and a US-only geographic profile. It is the first open-weight model on Bedrock to support explicit prompt caching, letting customers mark reusable prompt prefixes for reduced latency and discounted input-token pricing on cache hits.

CapabilityEconomicsAccessInfrastructure

Entities: Amazon Web Services, Amazon Bedrock, Moonshot AI, Kimi K3, Kimi K2, OpenCode

70Useful signal
1 source
1 primary
Was this useful?
01

What happened

AWS added Moonshot AI's Kimi K3 model to Amazon Bedrock, accessible via the bedrock-runtime Invoke/Converse APIs and OpenAI-compatible endpoints, with both a global cross-Region inference profile and a US-only profile. AWS states this is the first open-weight model on Bedrock to support explicit prompt caching, where customers mark reusable prompt prefixes to cut latency and get a discount (roughly 10% on the global profile) on cache-hit input tokens. This is a distribution and availability event: the model itself was already released by Moonshot AI; what changed is that Bedrock customers can now call it directly.

02

Why it matters

Developers and enterprises already using Bedrock get an easier way to run a large open-weight model for long-context and coding workloads without moving data outside their existing AWS security boundary, which is a real, near-term procurement decision for teams evaluating model options. The prompt-caching feature is concretely actionable now: it directly lowers cost and latency for repeated-prefix workloads like agentic coding tools (OpenCode, Hermes Agent are named as consumers). The broader competitive impact is limited: this does not change the standing of Moonshot AI versus DeepSeek, Qwen, or Mistral in the open-weight race, it just adds one more distribution channel.

03

What is noise

Moonshot AI's claims of 2.8 trillion parameters and "~2.5x scaling efficiency" over Kimi K2 come from the vendor and are passed through unverified by AWS and by this extraction. AWS's framing of "sustained investment" in open-weight choice is standard vendor packaging: Bedrock has added dozens of open-weight models before, so this is one more addition, not evidence of a strategic shift. The extraction also cites opencode.ai/config.json as primary evidence, which does not match the actual AWS blog source, a sourcing defect worth flagging.

04

Watch next

  1. 01Independent benchmark results for Kimi K3 (coding, long-context, agentic tasks) from a third party, not Moonshot AI or AWS, to check the 2.8T parameter and 2.5x efficiency claims
  2. 02Actual customer adoption signals on Bedrock (usage tier announcements, case studies, or pricing changes tied to Kimi K3) over the next 1-3 months, since availability alone does not mean uptake
  3. 03Whether other major clouds (Azure, Google Cloud, Oracle) or inference platforms add Kimi K3 with prompt caching, which would show this is a genuine model-level advance rather than an AWS distribution exclusive

Evidence

1 linked

Coverage

1 story

More capability signals

Full feed →