Microsoft shifts Copilot to usage billing and tests DeepSeek to cut AI costs
To combat soaring agentic compute costs, Microsoft is ending flat-rate subscriptions and exploring cheaper models like DeepSeek V4.
June 16, 2026

Microsoft is fundamentally reshaping the commercial landscape of enterprise artificial intelligence by moving its flagship agentic tool, Copilot Cowork, to a usage-based billing model. As the service transitions to general availability worldwide, the tech giant is departing from the flat-rate monthly subscriptions that have defined enterprise software for over two decades. In a parallel move aimed at curbing the compounding costs of running complex AI workflows, Microsoft is actively exploring the integration of a fine-tuned, self-hosted version of DeepSeek V4 as a lower-cost model alternative. Together, these strategies signal a critical industry shift, proving that the compute-intensive nature of autonomous AI agents has rendered traditional unlimited-seat software models unsustainable.
For years, Microsoft normalized flat-rate software pricing, packaging its standard Microsoft 365 Copilot as a thirty-dollar monthly add-on per user. However, the introduction of agentic tools like Copilot Cowork, which executes complex, long-running, multi-step workflows across business apps and data sources without constant human intervention, has fractured this pricing structure. Charles Lamanna, Microsoft’s executive vice president for Copilot, agents, and platform, bluntly acknowledged that flat-rate pricing is no longer tenable for highly active users who command hundreds of automated tasks per week. Comparing the new system to filling a vehicle's gas tank at the pump, Lamanna noted that heavy agent usage drives computing expenses extraordinarily high. To address this, Microsoft is introducing Copilot Credits, where enterprise customers are billed precisely for the computation their workflows consume, including model tokens, context retrieval, tool integration, and overall runtimes[1][2].
The technical reality of agentic AI explains why the unmetered seat model is failing. While a basic chatbot operates on a simple query-and-response loop, pausing immediately after delivering an answer, an agentic system like Copilot Cowork works dynamically in the background[3]. It must independently coordinate multi-tool tasks, query vast enterprise document libraries, verify intermediate results, and self-correct when errors arise[4][3]. This continuous planning and execution process consumes massive amounts of computing power and millions of data tokens[5]. To prevent corporate customers from suffering from bill shock, Microsoft is launching the service as disabled by default[1]. IT administrators will be given granular controls to set spending limits, monitor resource consumption across different departments, and configure system alerts, allowing companies to treat AI as a metered corporate utility rather than an infinite resource[1][2]. During its preview program, Microsoft observed that early users deployed Cowork for massive, resource-intensive tasks, such as comparing nearly four thousand files across product versions or analyzing stalled sales pipelines to rank at-risk opportunities[4]. These multi-day workflows, while highly valuable, require multiple computational steps, context-aware retrieval, and recurring API calls that exponentially increase the required processing power[1][6][4].
To offer enterprises a relief valve from high operational costs, Microsoft is exploring alternative, highly efficient model options[7][8]. While Copilot Cowork currently relies on premier models from Anthropic—such as Opus 4.8 and Sonnet 4.6—and offers OpenAI's GPT 5.5 for high-tier enterprise users, Microsoft is testing a fine-tuned version of DeepSeek V4[9][2][8]. Developed by the Hangzhou-based DeepSeek, this model series has gained widespread attention for delivering robust reasoning and code execution at a fraction of the cost of its Western competitors[9][10]. The financial incentive is hard to ignore, as soaring token fees from leading proprietary models have become a major barrier to widespread enterprise automation, with developers increasingly seeking alternatives to mitigate expensive agent loops[5][11]. By offering DeepSeek as an optional, lower-cost alternative alongside their planned, proprietary Cowork 1 model, Microsoft hopes to give budget-conscious enterprises a viable way to scale autonomous agents without risking financial strain[2][8]. The tech giant expects to officially introduce a low-cost model option in the coming weeks, providing much-needed flexibility to the platform[9][7].
The potential inclusion of a Chinese-developed AI model in Microsoft's premier business suite carries significant geopolitical weight, especially amid heightened scrutiny from regulators and politicians regarding technological supply chains[12][8]. Microsoft has anticipated these concerns by clarifying that any deployment of DeepSeek V4 would be fully hosted on Microsoft Azure infrastructure[7][13]. By hosting the model internally, Microsoft ensures that customer data remains strictly isolated within its own secure cloud environment, adhering to Azure’s rigorous enterprise security and compliance protocols[7]. This setup prevents customer data from ever leaving the company's secure boundary, striking a delicate balance between leveraging low-cost global innovations and maintaining national security and privacy standards[7].
Microsoft is far from the only technology company adjusting to the harsh financial realities of generative AI. The industry is witnessing a widespread transition to usage-based billing as other AI leaders hit identical economic walls[2]. Microsoft’s own developer subsidiary, GitHub, recently implemented consumption-based structures, sparking debate among programmers who saw their bills rise based on the volume of code generated[2]. Similarly, Anthropic announced that its latest cutting-edge models would shift to usage-based billing models, even for premium subscribers[2]. In the broader corporate landscape, stories of runaway AI budgets have become cautionary tales, with corporate IT administrators raising alarms over unexpected billing spikes caused by employee experimentation[5]. Some developers have even coined terms like tokenmaxxing to describe the phenomenon of employees running expansive, recursive agentic loops for relatively menial tasks[5]. This collective shift suggests that the era of subsidizing high-performance AI compute to gain market share is rapidly coming to an end, as providers and enterprises alike demand predictable, manageable cost structures[2][8].
Ultimately, Microsoft’s pivot to consumption-based billing and its pursuit of cheaper open-weights models like DeepSeek V4 represent a major milestone in the commercialization of artificial intelligence[7][13]. It marks the formal transition of generative AI from a novel, bundled feature into a metered infrastructure utility[3]. While this shift may initially frustrate enterprise buyers accustomed to predictable, flat-rate software budgets, it introduces a level of economic transparency that is vital for long-term scalability[8]. By forcing organizations to evaluate the direct return on investment for every automated task, Microsoft is setting a new precedent for how the tech industry will build, price, and consume agentic AI in the years to come[8][3].