The AI Wire

This Week’s AI Price Collapse Moves Your Job To Supervision

Kimi K3 and GPT-5.6's price cuts arrived this week, alongside a cheaper DeepSeek coding model. None of it makes you obsolete, but together it makes knowing which AI to direct the skill worth paying for.

This Week's AI Price War Makes Supervision The New Skill

This week’s AI news is about price, not new capability. Moonshot AI released a free 2.8 trillion parameter open-weight model, and OpenAI cut GPT-5.6 pricing by up to 80 percent. DeepSeek’s newest coding model separately priced input tokens at $0.14 per million. That price collapse moves your job from using AI to supervising it.

Takeaways

  • Moonshot AI released Kimi K3, a 2.8 trillion parameter open-weight model, on July 27, 2026, and Interconnects called it the largest open-weight model released to date.
  • OpenAI cut GPT-5.6 Luna pricing by 80 percent and Terra pricing by 20 percent on August 1, 2026, according to OpenAI’s own announcement.
  • DeepSeek’s V4-Flash-0731 exited beta on July 31, 2026, priced at $0.14 per million input tokens, with output token pricing listed separately.
  • Running Kimi K3 at full scale still requires a cluster of high-end GPUs, so free open weights do not yet let a solo builder self-host a frontier-class model on a laptop.
  • Choosing which model handles which task and checking its output is the skill worth building this week, not typing prompts faster.
Kimi K3 parameters 2.8 trillion
GPT-5.6 Luna price cut 80 percent
GPT-5.6 Terra price cut 20 percent
DeepSeek input price $0.14 per million tokens

What shipped this week?

Four announcements were published between July 27 and August 1, 2026, and three of them cut price rather than added capability. Moonshot AI released open weights for Kimi K3, a 2.8 trillion parameter model that Interconnects analyst Nathan Lambert called the largest open-weight model yet, with coding and reasoning scores close to closed frontier systems (Interconnects, July 27, 2026). Tom’s Hardware reported the same release runs roughly 2 to 3 times more efficiently than comparable frontier-scale models, and framed it as a shot across the bow of OpenAI and Anthropic (Tom’s Hardware, July 27, 2026). On August 1, 2026, OpenAI cut GPT-5.6 Luna pricing by 80 percent and Terra pricing by 20 percent (OpenAI, August 1, 2026). DeepSeek’s V4-Flash-0731 had exited beta the day before, priced at $0.14 per million input tokens (DeepSeek, July 31, 2026).

Why does cheaper AI change your job?

Cheaper AI changes your job because automation’s constraint used to be cost, and for most agent tasks it no longer is. When GPT-5.6 Luna drops 80 percent (OpenAI, August 1, 2026) and a capable coding model prices input tokens at $0.14 per million (DeepSeek, July 31, 2026), a manager no longer has to justify running an AI agent on a research task instead of assigning it to a person. The bottleneck moves to who decides which model handles which task, and who checks its output before it goes anywhere important. I’ve written before about how cheaper AI access gives workers leverage rather than just savings, because people who can direct several models well become harder to replace than people who can only use one (This Week’s AI Price Cuts Give Workers Leverage, Not Just Savings). Supervision is the skill this week rewards.

How big is Kimi K3, really?

Kimi K3 is a 2.8 trillion parameter open-weight model that Moonshot AI released in full on July 27, 2026 (Tom’s Hardware, July 27, 2026). Interconnects described it as the largest open-weight model released to date, with coding and agent benchmark scores close to closed frontier systems (Interconnects, July 27, 2026). Tom’s Hardware separately reported the model runs roughly 2 to 3 times more efficiently than comparably sized frontier-class systems (Tom’s Hardware, July 27, 2026). For a team blocked from sending code to an external API by data-privacy rules, this is the first time a self-hosted model has come close to matching what OpenAI or Anthropic sell as a service.

What’s the catch with ‘free’?

The catch is that open weights are not the same as accessible compute. A 2.8 trillion parameter model needs a cluster of high-end GPUs to run at usable speed, so a free download does not mean a solo builder can run Kimi K3 on a laptop (Tom’s Hardware, July 27, 2026). The claim that it performs close to frontier systems comes from Moonshot’s own release and Interconnects’ independent read of those benchmark numbers, not from a third-party audit across every task type (Interconnects, July 27, 2026). DeepSeek’s published $0.14 per million figure covers input tokens only. Output tokens, which dominate cost on long coding and agent runs, are priced separately and weren’t part of that headline number (DeepSeek, July 31, 2026). Treat near-free pricing as a starting point, not a total cost.

Six changes ranked by cost of ignoring

  1. Model selection becomes a weekly decision. When Moonshot, OpenAI and DeepSeek all move price or capability within a five day window (Interconnects, July 27, 2026), locking into one vendor for a year stops making sense.
  2. Self-hosting becomes a viable option for privacy-restricted teams. A model that runs 2 to 3 times more efficiently than comparable frontier systems (Tom’s Hardware, July 27, 2026) turns a permanent no into a question for IT.
  3. Budget stops being the blocker for agent projects. An 80 percent price cut on GPT-5.6 Luna (OpenAI, August 1, 2026) moves the conversation from affordability to trust.
  4. Hourly-billed junior tasks get absorbed into a token budget. Research and first-draft coding priced at $0.14 per million input tokens (DeepSeek, July 31, 2026) costs less than a minute of an analyst’s time.
  5. The gap widens between people who supervise AI and people who only prompt it. Cheaper access rewards judgment about which model fits which task, not typing speed.
  6. Vendors compete on price faster than on reliability. That race made agent output cheaper without making it more safely shippable unchecked, a pattern that played out this same cycle before (AI Agents Got Cheaper This Week, Not Safer).

Should you switch tools now?

You do not need to switch tools this week, but you should test one of these against a task you already trust. Run the same low-stakes research or coding task through GPT-5.6 at its new price and through DeepSeek’s V4-Flash-0731, then compare each against whichever model your team already uses for cost and output quality. If your organization blocks sending code to external APIs, ask whether IT would consider evaluating a self-hosted option like Kimi K3, since that conversation is now realistic in a way it wasn’t a year ago. Whatever you choose, keep a person reviewing the output before it reaches a client or a manager. I’ve written separately about how to explain AI-assisted work to a boss who doesn’t yet trust it, and that conversation gets easier when you can point to a specific model’s price and its documented limits (How to Explain AI-Assisted Work to a Boss Who Doesn’t Trust It).

Prompts you can use

Paste these straight in. Change the parts in square brackets and nothing else.

Compare Two AI Models
You are an AI infrastructure advisor helping me choose between two AI models for a specific work task. Before answering, ask me: what the task is (research, coding, drafting or analysis), how many tokens I expect to use per week, and whether my organization restricts sending data to third-party APIs. Once I answer, compare the two models I name on cost per week at my expected volume, whether either requires self-hosting, and what could go wrong if I trust the output without review. End with one specific recommendation, not a hedge.

Fill in the two models you’re actually choosing between and check current pricing yourself, since the assistant can’t verify live prices.

Build An Output Review Checklist
I used an AI model to complete a work task and need a checklist before I send the result to my manager or a client. Ask me first what the task was and what a wrong or misleading output would cost me if I missed it. Then give me a numbered checklist of the five most likely failure points for this specific kind of output, ordered by how costly each one would be to miss.

The checklist is only as good as your answer to its first question, so describe the task specifically rather than in general terms.

Explain An AI Tool Switch
Help me write a short explanation for my manager of why I want to switch from one AI tool to a cheaper alternative for a specific task. Ask me first what my manager's biggest concern about AI use is: cost, output quality, data privacy or job impact. Then write two paragraphs that name the actual price difference, address that specific concern directly, and avoid overselling the new tool beyond what I can personally verify.

Answer its clarifying question honestly before running it, or the result will read like generic marketing copy your manager will see through.

Questions people actually ask

Is Kimi K3 free to use?

The open weights are free to download under Moonshot AI’s license, but running the full 2.8 trillion parameter model requires a cluster of high-end GPUs, so a free download does not guarantee free or easy operation for most individuals or small teams.

How much cheaper did GPT-5.6 get?

OpenAI cut GPT-5.6 Luna pricing by 80 percent and Terra pricing by 20 percent on August 1, 2026. The cuts apply to API pricing, which mainly affects developers and businesses building on GPT-5.6 rather than consumer chat subscriptions.

What is DeepSeek V4-Flash-0731 good for?

DeepSeek positions V4-Flash-0731 as a coding and agent model, priced at $0.14 per million input tokens after exiting beta on July 31, 2026. That price covers input tokens only, and output token costs, which add up fast on long coding runs, are billed separately.

Can a small team self-host Kimi K3 without a data center?

Not easily. A 2.8 trillion parameter model needs a cluster of high-end GPUs to run at usable speed, so most small teams will still rent cloud GPU capacity rather than running Kimi K3 on office hardware, even though the weights themselves cost nothing to download.

Sources

  1. Interconnects, July 27, 2026interconnects.ai
  2. Tom’s Hardware, July 27, 2026tomshardware.com
  3. OpenAI, August 1, 2026openai.com
  4. DeepSeek, July 31, 2026huggingface.co

What happens next

Watch whether Anthropic or Google respond with their own price cuts in the coming weeks, since OpenAI and DeepSeek have both moved first this cycle. Also watch whether any independent benchmark outside Moonshot’s own release confirms Kimi K3’s frontier-adjacent scores, because right now the strongest claims trace back to the model’s own maker and one analyst’s read of its numbers.

Take this further

Full resume rewrite, section by sectionBest on Claude
Act as a blunt hiring manager who has read ten thousand resumes, not a career coach. I will paste my full resume and the job description I want. Rewrite the whole resume for that role, section by section, in this order: summary, experience, skills, education. Rules: every experience line leads with impact, not duty. Use bracketed placeholders like [8 percent] for any number I did not give you, and list at the end every placeholder I need to replace with a real figure. Keep it to one page of text. Plain formatting only, no tables or columns, so screening software can parse it. After the rewrite, tell me the three weakest claims that need evidence before I send this anywhere. My resume: [paste resume]. The role: [paste job description].
Share this
About the author
Vinayak Kapoor
Vinayak Kapoor

Vinayak started at seventeen on a call centre floor and climbed every rung himself over fifteen years: millions of customer conversations for some of the world's largest brands, self-taught design, video and web work, a profitable e-commerce brand of his own, and now human cyber risk, where he has built customer success journeys for national critical infrastructure and leads business growth at HumanFirewall. Nobody groomed him. He learned every skill alone, including the AI he now builds with daily as founder of Quarry, MaaSify and HuMatrix. He writes here so your career gets the guide his never had.

Read next
Come find me

The daily thinking
lives on the feeds.

Long form here. The working notes, the arguments and the things that did not fit go out most days.

← All writing