Microsoft Evaluating Kimi K3 Integration into Copilot to Reduce Inference Costs

Alina Collins
Published todayAbout 5 min read

Microsoft is evaluating Moonshot AI's open-source Kimi K3 for Copilot and Azure, driven by rising inference costs after Copilot shifted to per-token billing — but the model's deployment costs rival Claude Sonnet's, so actual savings remain unproven.

01

Why is Microsoft eyeing a Chinese open-source model?

Copilot Cowork has switched to per-token billing — every AI inference now carries a direct cost line item, sharpening Microsoft's cost sensitivity.
In June, Microsoft evaluated whether DeepSeek V4 could power Copilot Cowork. Testing Kimi K3 follows the same logic — finding cheaper alternatives to OpenAI and Anthropic.
This means → Microsoft is not betting on one model. It is systematically scanning the market for anything that can cut the inference bill.
02

How capable is Kimi K3?

Kimi K3 is Moonshot AI's latest model: 2.8 trillion parameters, multimodal, with an input context window of roughly 1 million tokens.
On the Arena coding leaderboard, it reportedly outperformed OpenAI's GPT-5.6 and Anthropic's Claude Fable 5.
In plain terms = by benchmark scores, the model is strong enough. Microsoft is not settling — it is looking for "good *and* cheap."
03

Does a cheaper model actually save money?

Arena noted that Kimi K3's 2.8 trillion parameters make it expensive to run. Deployment costs are comparable to Anthropic's Claude Sonnet series.
This means → open-source access and a low entry barrier do not equal low operating costs. The real expense is in deployment, not in obtaining the model.
Microsoft's engineering team is still assessing whether Kimi K3 can be deployed in Copilot. Cost savings depend on engineering optimization, and no conclusion has been reached.

Content is for reference only, not financial advice.

Microsoft Evaluating Kimi K3 Integration into Copilot to Reduce Inference Costs · nashnova