Skip to content

DeepSeek has brought springtime for AI application entrepreneurs

Aug 1, 11:42
Original Title: "DeepSeek Brings the Spring of AI Application Entrepreneurs"


On July 31st, DeepSeek V4-Flash officially launched its API for public testing.


For every million input tokens, it costs $1; for every million output tokens, it costs $2. After the cache is hit, the cost for every million input tokens is only 2 cents. It supports 1 million contexts, thinking modes, tool calls, and Responses API, with a maximum of 2500 concurrent connections per account.


Next, DeepSeek will adopt surge pricing, doubling the price during peak hours. Nevertheless, the price is still quite affordable.


According to a rough conversion provided by DeepSeek, approximately 0.6 tokens are consumed by one Chinese character. With one million tokens, you can read about 1.67 million Chinese characters. Based solely on input costs, reading about 1.67 million Chinese characters will only cost one dollar.


Over the past year, the large model industry has become accustomed to longer contexts, more complex reasoning, and increasingly capable agents. With each model update, discussions often revolve around who has surpassed whom, how much the benchmark has moved forward, and how far we are from AGI.


For AI application entrepreneurs, there is a more practical question:


How much will a task actually cost.


The model's capabilities determine what the product theoretically can do, while the pricing of calls determines how much of that capability the team dares to truly embed in the product.


This afternoon, I had a long conversation with ChatGPT via voice. I said, the current business model for AI applications that work is essentially token arbitrage.


Several times, it tried to help me phrase it more decently, such as "retail reasoning capabilities" or "intelligent service packaging." I said, don't bother finding a more polite term for this business, arbitrage is arbitrage, being blunt doesn't mean it's wrong.


AI applications seem to be the newest entrepreneurial direction in recent years, but at its core, it is a very old business:


Procurement, processing, pricing, and then selling the goods.


And many industrial changes have started from a sudden drop in procurement prices.


AI Applications are Essentially a Token Arbitrage Business


In the past, people were used to understanding AI applications within the framework of Internet products or SaaS.


After software development is completed, it can be distributed to a thousand people or a million people. Adding a new user naturally incurs bandwidth, storage, and service costs, but these costs are usually gradually diluted as the scale increases.


AI applications are different. Every time a user generates an image, writes a report, completes a piece of code, or has the Agent perform an additional step, the application company needs to re-invoke the model. Every time a user clicks a button, the upstream entity gets paid again.


Therefore, the basic business model that most pure software AI applications currently run on is actually quite simple:


Procure tokens from a model company, package the general model capability into writing, programming, design, customer service, video, or Agent products, and then sell them to users through memberships, points, and pay-per-use fees.


Take Poe as an example. When launching its API, it directly stated that the pricing goal for additional points is to charge based on the price the underlying model supplier charges it.


The upstream entities charge for the model and tokens, Poe converts them into points on the platform, and then sells them to users along with a unified entry point, payment method, model switching, and development interface.



Some companies may say they do not sell tokens, they sell outcomes. The customer service software company Intercom designed a price for AI customer service Fin, charging $0.99 for each completed outcome. When a user confirms the issue is resolved, does not require further human assistance, or when Fin successfully completes a workflow, it can be considered a completed outcome.


It's just a different pricing label — charging per seat, per task, or per successfully resolved issue — the change is the pricing unit the user sees. Behind the scenes, the application company still needs to invoke the model and calculate how many tokens are needed to complete a delivery.


Southern Song Dynasty painter Li Song once painted a piece called "Market Vendor Playing with a Child on a Shoulder Pole." The vendor carried a load of goods as he walked through villages and alleys, with pots, bowls, needles, snacks, and toys all hanging on the shoulder pole. A load of goods was like a mobile convenience store, and the vendor earned money from selecting, packaging, and delivering goods.


AI application companies do something similar. The model manufacturer produces intelligence, and the application company applies that intelligence to specific scenarios. What is bought at the lower level is a token, and what should ultimately be sold is a report, a design drawing, or a video.


The model company sells intelligence, and the application company sells outcomes.


The profit of this business mainly depends on three price differences.


The first one is the purchase price difference.


For the same model, the actual cost obtained by different teams may not be the same. Companies with a high volume of requests can negotiate discounts, while teams with strong engineering capabilities can utilize caching, asynchronous processing, peak/off-peak call patterns, and model routing to further reduce the average cost. What others can do for $1, you can do for $0.50, and the remaining $0.50 is your profit.


In simpler terms, when purchasing the same item, some may get it at wholesale prices, while others may get it at retail prices.


The second one is the usage price difference.


Some people, after becoming members, use AI intensively every day, while others, full of ambition when recharging, stop using the product after just two weeks. As long as the average consumption of the entire user base is below the estimated level when designing the package, the unused quota becomes a profit margin.


This logic is similar to that of gym memberships and mobile data plans.


However, there is one thing to note here. The upstream purchases need to be flexible enough so that more profit is generated when users use the service less. If a company has already signed a fixed yearly commitment and the remaining tokens are just sitting in the account, that is not profit; it is unsold inventory.


Unused portions by users are called stagnant assets, while unsold inventory by the company is called overstock.


The third one is the product price difference.


Consuming the same ten thousand tokens, some products only output a piece of text that requires extensive user editing, while others have completed data organization, content generation, fact-checking, format adjustments, and final export. The underlying costs may be similar, but the prices that users are willing to pay can vary significantly.


The true value of an AI application depends on this. Purchasing determines whether this business can survive, and the product determines how much this business is worth.


The resale of tokens is just the foundation; the product is the reason for the price difference.


Product Managers Have Become Accountants First


The trouble lies in the fact that an AI application needs constant purchases, so product issues easily become cost issues.


In Manus's official usage instructions, you can already see how this cost shapes the product. Manus deducts points from the monthly subscription based on the complexity of the task and the resources needed. When asking AI questions, for simple queries, the official recommendation is to use the Chat mode as much as possible, without initiating the full-fledged autonomous agent; for complex tasks, it's best to check intermediate results first to prevent the agent from continuing in the wrong direction and consuming more points unnecessarily.



This advice is certainly reasonable. However, it also reveals the most troublesome aspect of Agent products: the more proactive the model, the harder it is for the product company to predict costs.


A chatbot stops after answering a question once, making it relatively easy to calculate costs. An Agent, on the other hand, will autonomously break down tasks, search repeatedly, open web pages, call upon tools, and then decide on the next steps based on interim results. While the user may have issued only one command, behind the scenes, dozens of model calls may have already occurred.


As a result, products have started to teach users how to minimize the workload on the Agent. The fear is that after a long period of number crunching, the accountant will start taking over the product. When considering whether to launch a new feature, the team's primary concern is no longer whether the user needs it, but rather how expensive each individual task will be. While the model has acquired greater capabilities, the application is hesitant to fully entrust these capabilities to the user.


Of course, cost control is essential for a company. No startup can afford to treat the most expensive model as a commodity for users to freely use.


Some products have placed a more effective model in a high-priced package, offering ordinary users a version that has passed through multiple layers of throttling. The model's capabilities have not disappeared; they have simply been allocated to different users after the cost calculation.


Over time, the industry has witnessed a rather awkward inversion.


Users purchase AI applications hoping the model will understand more, think further, and complete additional tasks. However, to maintain costs, application companies must limit the model's reading, thinking, and actions.


While the product is marketed as intelligent, the team's most practiced action has become to withhold intelligence.


Furthermore, if prices rise upstream, application companies must reduce their allocations; if competitors obtain lower costs, membership fees in the market will immediately decrease. When new inexpensive models emerge, existing routing and packages need to be recalculated. Application companies may appear to control the product, but the most critical variable is always held by upstream providers.


Over the past two years, AI application entrepreneurs have not necessarily lacked ideas. What they truly lack is intelligence that is strong and affordable enough to be confidently integrated into products.


DeepSeek Eliminates the "Toll of Intelligence"


The value of DeepSeek lies in the fact that it is not merely a model with reduced capabilities created to lower costs. It has around 1 million contexts, supports thinking modes, and tool invocation.


When DeepSeek released its preview version in April, it claimed its inference capabilities were close to V4-Pro, performing similarly in simple Agent tasks. The official version released on July 31st did not alter the model's structure and scale significantly. It mainly underwent retraining, and several Agent benchmark scores have now surpassed those of the V4-Pro preview version.



This is important.


Cheap models have existed before. The problem is, getting something cheap doesn't mean it's always subpar. What AI applications truly need is intelligence that is as stable as possible below a certain unit price.


At least from the benchmarks of Agent and code published by DeepSeek, V4-Flash has already crossed the "sufficient" line for a batch of high-frequency tasks, with a price low enough for the team to confidently make high-frequency calls.


For most ordinary users, large models have likely reached a tipping point.


The average person doesn't care about benchmarks. They rarely pay attention to whether a model is three or five points ahead in a certain test; they only care about whether their own tasks can be done effectively. When a model can already read materials, conduct reasoning, use tools, and deliver results, increasing its capabilities further is certainly valuable. But can this added value still justify a several-fold price increase? That needs to be recalculated.


In the past, model companies claiming to be stronger and more expensive seemed reasonable. However, this rationale does not hold up well now. It is necessary to discuss where the added cost comes from, what problems the additional capabilities solve, and how many users actually need them.


DeepSeek did not make all models cheaper; it made high-priced models into something that needs to be explained to users.


What Pinduoduo truly broke through in its early days was not just the prices of goods. It allowed consumers to see for the first time that many daily necessities could be sold so cheaply. The brand premium, channel costs, and intermediary links that were previously taken for granted suddenly needed some justification.


What Pinduoduo broke through was the price perception of goods.


What DeepSeek broke through was the price perception of intelligence.


Entrepreneurs Can Finally Focus on Product Development


For the AI application industry, the most important thing next is that entrepreneurs can finally shift their focus back from upstream.


In the past, rapid updates to large models often led application companies astray. Today integrating longer contexts, tomorrow switching to new reasoning models, and the day after showcasing Agent capabilities on the homepage. The product always seems to be upgrading, but the roadmap is not set by themselves. Whatever the model manufacturers release, application companies demonstrate.


A typical scenario of upstream prosperity and downstream distraction.


Now, foundational models have the opportunity to become a relatively stable production material. Teams no longer need to revolve around model rankings every day or consider "integrating the latest models" as the most important product progress.


Once the model has stabilized, the product finally runs out of excuses.


While the model's ability to perform a task is just the beginning, the product's job is to integrate the task into a real process, dealing with permissions, data, collaboration, reviews, and failures. This way, users don't have to reteach the AI how to work each time. These aspects are not as glamorous as model parameters and are challenging to benchmark. However, whether users are willing to continue paying in the second month often depends on these details.


The value of an AI application lies not in showcasing the model to users but in making the model disappear into the workflow.


When users no longer need to choose a model, delve into research prompts, or worry about the number of inference rounds behind a task, the application has truly done its job.


At this point, what sets application companies apart in competition is no longer who gets the model first but who understands a task better and who can transform the industry's uncodified rules into part of the product.


Models can become stronger overnight, but industry understanding cannot. APIs can be accessed by everyone, but a set of work processes truly accepted by users is challenging to replicate.


This shift has given many seemingly niche businesses a chance to thrive once again.


AI startups don't necessarily have to start from the "next-generation entry point," nor does every company have to create a super agent to serve all of humanity. AI can simply solve one problem in international trade inquiries, streamline a set of processes for chain stores, or eliminate a piece of repetitive work for a specific group of professionals.


The market doesn't have to accommodate everyone. As long as the problem is specific enough, the delivery is stable enough, and the value users receive exceeds the price they pay, the business can survive.


In the past, when models were too expensive, small markets struggled to sustain a complete product. Teams had to aim for a larger user base, higher average revenue per user, and faster funding to cover the ongoing inference costs. Many startups began telling overly grand stories without really finding their true users.


Now, the story can be smaller.


A small team can start by serving a small group of people, fully mastering one thing, and then gradually expanding.


This resembles the early days of many "century-old shops" today. They didn't aim to dominate a whole street when they first opened; they just focused on welcoming a few tables of customers at a time. With good food, reasonable prices, and repeat customers, the business naturally grew.


While the big model industry talks about scale, parameters, and endgames, application startups ultimately need to return to this simple logic.


DeepSeek does not answer the three questions for entrepreneurs: whom do you serve, what problem do you solve for them, and why can't they find a replacement.


It simply frees the team from spending most of their time negotiating with upstream partners. It also allows small teams that lack the ability to train large models and do not qualify for special pricing to get a relatively fair entry ticket.


In the past, the model itself was a barrier. Next, everyone can buy the model, but the difference lies in what they do with it. The truly scarce resources will gradually shift towards users, data, processes, and trust. Those closer to real work will have more opportunities.


After DeepSeek has lowered the threshold of intelligence, AI applications have finally moved from a model competition to a product business.



In his essay "On the Costliness of Coarse Grains," Han Dynasty's Chao Cuo wrote: "Coarse grains are valuable, but gold and jade are humble."


In recent years, large models have been more like gold and jade. The parameters have grown larger, the rankings have soared higher, placed under the spotlight at conferences, everyone can see their value, but few can truly incorporate them into daily work.


What the application layer needs to do is turn gold and jade into coarse grains.


DeepSeek has not made the meal for entrepreneurs; it has only lowered the price of rice.


As for who can open the restaurant, it depends on their respective skills.


-END-


Original Article Link


Recommended

Opinion: Why is a 7,709x Leveraged ETF Essentially a Natural Negative EV Product?

Jul 31, 17:30
Opinion: Why is a 7,709x Leveraged ETF Essentially a Natural Negative EV Product?

In July, only 153 venture capital firms made investments, is the crypto venture capital industry experiencing a "mass extinction"?

Jul 31, 16:20
In July, only 153 venture capital firms made investments, is the crypto venture capital industry experiencing a "mass extinction"?

Mainstream Media Reviews AI Stock Market Wizard's $45 Billion Liquidation: Silicon Valley's "Genius Worship" Collapse

Jul 31, 15:58
Mainstream Media Reviews AI Stock Market Wizard's $45 Billion Liquidation: Silicon Valley's "Genius Worship" Collapse

The Truth Behind the Fall of the 25-Year-Old AI Stock Market Prodigy: Not Just a Leverage Death Story

Jul 31, 15:52
The Truth Behind the Fall of the 25-Year-Old AI Stock Market Prodigy: Not Just a Leverage Death Story

CL, a Trading-Savvy Cat | Meet the Hyperliquid Trader

Jul 31, 15:33
CL, a Trading-Savvy Cat | Meet the Hyperliquid Trader

Goldman Sachs Raises Microsoft Price Target to $640, Betting on Surge in Copilot Subscriptions

Jul 31, 14:41
Goldman Sachs Raises Microsoft Price Target to $640, Betting on Surge in Copilot Subscriptions