Four separate AI releases landed inside seven days, and together they drop the cost of testing an agent on a task from your job close to zero. Moonshot AI open-sourced a near-frontier model, OpenAI cut its flagship pricing, DeepSeek shipped a cheap coding agent, and Perplexity brought a desktop agent to Windows.
Takeaways
- Moonshot AI released open weights for Kimi K3, a 2.8 trillion parameter model, on July 27, 2026, but self-hosting it requires GPU infrastructure most individuals don’t have.
- OpenAI cut GPT-5.6 Luna pricing by 80 percent and Terra pricing by 20 percent on August 1, 2026, making high-volume agent testing affordable for solo builders.
- DeepSeek’s V4-Flash-0731 exited beta on July 31, 2026 priced at $0.14 per million input tokens, giving professionals a near-free option for testing agent coding tasks.
- Perplexity’s desktop agent reached Windows on July 28, 2026 and can operate Word, Excel and Outlook directly, but it’s locked behind a $200-a-month Max plan.
- Pick one manual task and test it against DeepSeek or GPT-5.6 Luna this week, since both are now priced low enough to fail cheaply.
| Kimi K3 size | 2.8 trillion parameters |
|---|---|
| GPT-5.6 Luna cut | 80% price cut |
| DeepSeek V4-Flash price | $0.14 per million input tokens |
| Perplexity Windows agent | $200/month (Max plan) |
What changed for you this week?
Four separate AI releases landed inside seven days, and each one lowers the price of trying an AI agent on your job. Moonshot AI published open weights for Kimi K3, a 2.8 trillion parameter model that Tom’s Hardware reported on July 27, 2026 performs almost as well as frontier models. OpenAI cut API pricing for GPT-5.6 on August 1, 2026, dropping the Luna tier by 80 percent and the Terra tier by 20 percent, according to OpenAI’s own announcement. DeepSeek exited beta with V4-Flash-0731, a coding and agent model priced at $0.14 per million input tokens as of July 31, 2026, per its Hugging Face model card. And Perplexity brought its desktop agent to Windows on July 28, 2026, per SiliconANGLE, letting it operate Word, Excel and Outlook directly on your machine. None of these four are the same product, but together they make testing an agent on one task from your job cost close to nothing this week.
Can you self-host Kimi K3?
Not on your laptop, and that’s the caveat worth naming before you get excited about a free frontier model. Kimi K3 has 2.8 trillion parameters, and even at what Moonshot AI describes as two to three times easier to run than comparable frontier models, per Tom’s Hardware’s July 27, 2026 report, that vendor claim still describes GPU infrastructure costing tens of thousands of dollars, not a laptop or a single server. Self-hosting only makes sense for a company with data-privacy limits serious enough to justify that hardware instead of paying per token. If you work somewhere a compliance team blocks API calls to outside models, this is the release to bring to your infrastructure lead. If you’re an individual professional trying to save money, Kimi K3 is not your answer this week; the two API-priced models below are. I cover the wider gap between a vendor claim and what a policy proves in AI Agents Got Cheaper This Week, Not Safer, and it applies here too.
Why did GPT-5.6 get so much cheaper?
OpenAI cut GPT-5.6 pricing on August 1, 2026 because the fastest way to keep developers building on an API is to make the cheapest tier hard to ignore, and an 80 percent cut to the Luna tier does exactly that, according to OpenAI’s pricing announcement. The Terra tier, aimed at harder tasks, dropped a smaller 20 percent, per the same OpenAI post. For a solo builder or a small team running agent workflows, that Luna cut turns a task that cost meaningful money in July into one you can run many times over for the same budget in August. I wrote about the leverage this kind of repeated price cut hands workers, not just budgets, in This Week’s AI Price Cuts Give Workers Leverage, Not Just Savings. The pattern holds again here: cheaper tokens mean you can test five ideas instead of one before committing work time to any of them.
What does a $0.14 agent model change?
DeepSeek’s V4-Flash-0731 exited beta on July 31, 2026 priced at $0.14 per million input tokens, and at that price a full day of testing an automation idea costs cents, not dollars, per DeepSeek’s Hugging Face listing. The model card reports gains specifically in agent coding tasks, the kind of write-code, run-it, read-the-error, fix-it loop that used to require a subscription-tier model to do reliably. That combination, near-free and agent-capable, is what makes this week different from an ordinary price cut. It’s the first time a model this cheap has been positioned specifically for the kind of task-running professionals want to hand off. Treat the vendor’s benchmark numbers as a starting claim, not a guarantee, and test it against your own task before trusting it with anything sensitive.
Four ways to test it, ranked by cost
Start with this list, ordered from cheapest to most expensive, if you want to test one of these tools before the week ends.
- DeepSeek V4-Flash-0731 for a coding or research task. At $0.14 per million input tokens, per its Hugging Face card, this is the cheapest way to see what an agent model does with a multi-step task this week.
- GPT-5.6 Luna for a high-volume task. The 80 percent price cut, per OpenAI, makes it worth testing on anything you’d normally run many times, like drafting variations or summarizing a long document set.
- Perplexity’s Windows agent for one Office task. At $200 a month for Max users, per SiliconANGLE, this only pays for itself if you spend several hours a week moving data between Word, Excel and Outlook.
- Kimi K3 as an infrastructure proposal. With 2.8 trillion parameters, per Tom’s Hardware, this is a conversation to start with an IT or security team, not something to run yourself this week.
Is the Perplexity Windows agent worth it?
For most people, not yet, and the price is the reason. Perplexity’s desktop agent reached Windows on July 28, 2026, per SiliconANGLE’s report, and it can operate inside Word, Excel and Outlook directly rather than through copy-paste or a browser tab. That’s a genuine step past the browser-only computer-use agents most people have tried, and I explain how that category of tool works, and where it still breaks, in Long Context, Tool Use And Computer Use Explained Plainly. But it’s locked behind Perplexity’s $200 a month Max plan, per the same SiliconANGLE piece, which puts it well past casual use. It’s worth testing only if you can point to a specific weekly task, like reconciling a spreadsheet against an inbox, that currently costs you more than a few hours a month. If you can’t name that task, the $200 is paying for curiosity, not productivity.
What should you do this week?
Pick one task you already do by hand and run it through one of these tools before the week ends, because the point of this batch of releases is that testing no longer costs enough to justify waiting. Start with DeepSeek or the GPT-5.6 Luna tier since both are priced low enough to fail cheaply, per DeepSeek’s release notes and OpenAI’s pricing page. If the task involves data your company won’t let leave the building, that’s the moment to raise Kimi K3 with whoever owns your infrastructure budget rather than trying to solve it yourself. None of these four tools guarantees good output on your specific task; the price cuts only guarantee that finding out is now cheap.
Prompts you can use
Paste these straight in. Change the parts in square brackets and nothing else.
You are an automation consultant helping me evaluate whether an AI agent can reliably handle one specific task from my job. I will describe the task in detail, including the inputs I start with, the steps I currently take by hand, and what a correct finished output looks like. Your job: (1) tell me honestly whether this task is a good fit for an AI agent right now, or whether it needs a human in the loop at some step, (2) list the specific failure points I should check the output against before trusting it, (3) suggest how I'd verify the output is correct without redoing the whole task myself. Do not assume the task will work well just because it sounds automatable. If you need more detail about my task before answering, ask me first. My task is: [describe your task, the tools or data involved, and roughly how long it takes you by hand].
Works with any model, including the cheaper DeepSeek or GPT-5.6 tiers; the value is in the honest gap analysis, not the model brand.
You are a pragmatic technology advisor helping me decide whether my team should keep using an API-based AI model or consider a self-hosted open-weight model like Kimi K3. Ask me about: the sensitivity of the data involved, our current monthly spend on AI API calls, whether we have any in-house infrastructure or DevOps capacity, and our compliance or data-residency requirements. Based on my answers, give me a clear recommendation with the main tradeoff spelled out (cost, control, and maintenance burden), not a vague 'it depends.' If my answers suggest self-hosting isn't realistic for us yet, say so plainly and tell me what would need to change for that to make sense.
This is a decision-framing prompt, not a technical migration plan; loop in your actual infrastructure or security lead before acting on the output.
Act as a workflow auditor. I'm going to describe a recurring task I do that involves moving information between two or more of these: email, a spreadsheet, a document, or a chat tool. Ask me clarifying questions about the exact steps, how often I do it, and how much time it takes each time. Then tell me: (1) which parts of the task are mechanical enough to hand to an AI agent, (2) which parts need my judgment and should stay manual, (3) a rough estimate of whether the time saved is worth the cost of a tool priced around $200 a month, based on how often I do this task. Be specific and skeptical, not optimistic by default.
Useful before paying for a desktop agent like Perplexity’s; swap in your own numbers rather than trusting the model’s time estimates blindly.
Questions people actually ask
Can I run Kimi K3 on my own computer?
No. Kimi K3 has 2.8 trillion parameters, and even though Moonshot AI says it’s two to three times easier to run than comparable frontier models, per Tom’s Hardware’s July 27, 2026 report, that still means dedicated GPU infrastructure, not a laptop. It’s a decision for a company’s infrastructure team, not an individual test.
Is DeepSeek V4-Flash safe to use for work tasks?
DeepSeek’s own July 31, 2026 model card reports gains in agent coding tasks at $0.14 per million input tokens, but that’s a vendor benchmark, not an audit of your specific data-handling needs. Test it on a low-stakes task first, and check your company’s policy on sending data to outside models before using it on anything sensitive.
What’s the difference between GPT-5.6 Luna and Terra?
Per OpenAI’s August 1, 2026 announcement, Luna is the cheaper tier and got an 80 percent price cut, while Terra is aimed at more demanding tasks and got a smaller 20 percent cut. Luna is the one worth testing first if you’re running high-volume, repetitive tasks.
Do I need Perplexity Max to use the Windows agent?
Yes. Perplexity’s desktop agent on Windows, reported by SiliconANGLE on July 28, 2026, is available to Max subscribers at $200 a month. It’s only worth that price if you can point to a specific weekly task inside Word, Excel or Outlook that currently eats several hours.
Sources
- Tom’s Hardware reported on July 27, 2026tomshardware.com
- OpenAI’s own announcementopenai.com
- Hugging Face model cardhuggingface.co
- SiliconANGLEsiliconangle.com
What happens next
Watch whether other labs match DeepSeek’s per-token pricing on agent-specific models, since that number determines how cheap testing stays. Also watch whether Kimi K3 self-hosting shows up in enterprise procurement conversations by the end of 2026, which would confirm the privacy-driven demand Tom’s Hardware described rather than just a headline claim.
Take this further
Act as a blunt hiring manager who has read ten thousand resumes, not a career coach. I will paste my full resume and the job description I want. Rewrite the whole resume for that role, section by section, in this order: summary, experience, skills, education. Rules: every experience line leads with impact, not duty. Use bracketed placeholders like [8 percent] for any number I did not give you, and list at the end every placeholder I need to replace with a real figure. Keep it to one page of text. Plain formatting only, no tables or columns, so screening software can parse it. After the rewrite, tell me the three weakest claims that need evidence before I send this anywhere. My resume: [paste resume]. The role: [paste job description].

