Skip to content

The Silicon Valley That Copies China's Homework, Next Question: AI Applications

Aug 4, 11:36
The Silicon Valley That Copies China's Homework, Next Question: AI Applications
Article Title: "Silicon Valley Copying China's Homework, Next Question Is AI Applications"
Article Author: TrendWatch Beating


Silicon Valley has started copying China's homework.


Recently, the U.S. has come up with a batch of low-cost open-source models, aiming to catch up with China's open-source models in terms of capabilities while pricing them at the same level.


Mira Murati's Thinking Machines Lab released the first model, called Inkling, last month. The official statement mentioned two things: the underlying architecture was based on DeepSeek-V3, and the training data was generated by Kimi K2.5.


The most expensive people in Silicon Valley are using the framework of Chinese open-source models and feeding them with data generated by Chinese models. If this had happened two years ago, it would have been the other way around.


Earlier this year, Arcee AI's Trinity-Large-Thinking scored 91.9 on PinchBench, 1.4 points lower than Anthropic's Opus 4.6, at 4% of the price of the latter.


Arcee AI's CEO Mark McQuade said, "We are working hard to help the U.S. catch up with and surpass China."


Now, in Silicon Valley's copying endeavor, it's not just the models being replicated but also the prices set by Chinese companies.


Architectures can be referenced, training methods can be reproduced, and data can be handed over to other models for batch generation. Once the weights are released, developers can continue to build on them. As for the price, as long as one company starts cutting prices, the rest will sooner or later follow suit.


Over the past three years, the big model companies have built a towering pagoda based on parameters, computing power, and closed sourcing, entertaining guests. Now, this building is on the brink of collapse.


Sixteen Days


On July 17th, in the early morning, the dark side of the moon released Kimi K3. With a total of 2.8 trillion parameters, 1 million tokens of context, it is the world's largest open model.


In the Frontend Code Arena programming leaderboard, K3 scored 1679, ranking ahead of Claude Fable 5 and GPT-5.6 Sol. This is the first time an open model has surpassed closed-source models. In the seven frontend sub-segments, K3 took the first place in six of them.


Under the evaluation results, Musk left a comment, "Impressive."



K3's API is not cheap, charging $100 per million output tokens. However, in the SuperCLUE's horizontal test, its task score was 18% higher than the second place, and after running the same batch of tasks, the average cost actually decreased by 16%.


Whether a model is expensive or not should not be judged by the price per million tokens, but by how much it costs to complete a task from start to finish. The further apart these two numbers are, the more meaningless the price per token is.


In the late night two days later, the dark side of the moon paused C-end new user subscriptions. The request volume in 48 hours far exceeded the estimate, and the computing power was insufficient.


On July 27, K3 opened up its full weight, along with the technical report and supporting infrastructure.


Next up was DeepSeek.


The market had been waiting for the official release of V4. At the end of April, DeepSeek first released two preview versions, V4-Pro and V4-Flash, both supporting 1 million token contexts. The official term for this was "the era of million-token contextual inclusion." There was little news in the next two to three months until August 2, when V4 Flash was officially released with open weight. With a total of 284 billion parameters, only 130 billion were activated in a single inference, and the DeepSWE programming test score rose from 7.3 to 54.4.


What was even more aggressive was the pricing. In May, V4-Pro's price was directly reduced to one-fourth of the original price, marking the fourth price adjustment by DeepSeek in a month. When Flash officially launched, everyone was still surprised by its cost-effectiveness.


"You get what you pay for" is a thing of the past. Now Chinese models are becoming more competitive while charging less. Foreign competitors are naturally uncomfortable, but there's not much they can do. Users have seen something cheap and easy to use, and no one wants to go back to being a sucker.


Following that, Alibaba announced that it would release the weight for Qwen3.8-Max next week. The Max series, which has always been in a closed-source flagship position, has also begun to open up. Its international price is approximately 40% of Opus 5's input cost and only 24% of its output cost.


All these events took place within a mere 16 days.


In the past, model companies used scarcity to set prices, saying, "Only I can do it, so you have to pay my price." With just a few points difference on the leaderboard, prices could differ by tens of times.


The technical roadmaps of Kimi, DeepSeek, and Qwen are different, but they are all working towards the same goal: pushing capabilities up, opening up weights, lowering prices, and quickly incorporating them into products.


Wanting to Copy Homework, But No One Is Paying


Some people in the United States have also understood this trend.


By the end of 2025, Arcee AI almost fully committed the money in its account. In just 33 days, with 2048 B300 units and $20 million, the entire company had only 30 people. The Trinity Large produced has a total parameter of 400 billion, with a single inference activation of 13 billion, and nearly half of the training data is synthetic. In July of this year, Arcee also signed a collaboration with the U.S. Department of Energy, preparing to develop a model for scientific research.


The team, achievements, costs, government orders—everything is in place. However, when it comes to financing, they just can't seem to get through.


To date, Arcee has raised a total of $50 million, with a valuation of $240 million. In today's world of large models, this amount of money is not even enough to reserve a table, let alone access the private room.


CEO Mark McQuade said, "Almost all top-tier VCs have rejected us."


The open-source model from the U.S. has not yet directly competed with China and was first stopped by its own investment committee.



In the first quarter of 2026, global AI startups raised $255.5 billion, with close to two-thirds concentrated in three transactions: $122 billion for OpenAI, $30 billion for Anthropic, and $20 billion for xAI. Most of the money flowed into closed-source companies.


Joe Floyd, an early investor in Arcee and a partner at Emergence Capital, said that the reasons for rejection he heard from many VCs were similar:


"I don't want this model to succeed. I don't want to invest because this will harm my investments in Anthropic and OpenAI."


Ironically, those willing to spend money on the open-source model are chip sellers.


In March of this year, NVIDIA disclosed in SEC filings that it plans to invest $26 billion over the next five years to support open weights. Reflection AI, Poolside, and Thinking Machines Lab have all received funding from them.


This bill was the one Huang Renxun understood better than anyone else. The closed-source model makes money from API, while the open-source model burns through GPUs. The cheaper the model is sold, the more developers jump onboard, and NVIDIA's GPUs sell even better.


On the evening of July 24, Huang Renxun registered an X account and tweeted for the first time by forwarding an open letter titled "Open Weight and America's Leadership in the AI Field." The letter mentioned that for the United States to continue leading this industry, it should not only focus on having the most advanced models but also on having an ecosystem that can spread across various industries.


Later, nearly 200 American startup founders jointly wrote to the Trump administration opposing the ban on Chinese open-source models. After reading this letter, the most interesting part was that there was no catchy phrase. They did not speak positively about China or advocate for openness. Their only reason was that affordable Chinese models had already become a production tool. If these were banned, they would have to go back and buy expensive American APIs.


McQuade also mentioned another statement during an interview, saying that mass adoption was only a matter of time.


Demand Driven by Affordability


While Silicon Valley is still debating whether to adopt open weights, China has already moved on to the next challenge.


The phrase "The Year of AI Applications" has been echoed repeatedly over the past three years. However, determining whether a technology has truly entered the application stage depends only on two things: whether users are willing to pay and whether enterprises have integrated it into their operations.


This year, the annual revenues of Claude Code and Cursor have both exceeded $1 billion, and the penetration rate of AI programming tools among individual developers is close to 50%.


The transformation on the enterprise side is even more significant. According to a study by Sand Dune Think Tank, the adoption rate of AI agents by Chinese enterprises increased from 17.3% at the end of 2024 to 40.3% by mid-2026.


For this, Li Yanhong proposed a new metric called DAA, Daily Active AI, which measures how many AI agents are actively working each day and how many tangible results they are delivering, shifting the focus from the number of daily product users in internet companies to the labor done by AI agents.


The surge in API calls is even more astonishing. Following CCTV's calculation method, China's daily token call volume increased from about 1 trillion at the beginning of 2024 to 140 quadrillion by the end of March this year, a growth of over 1,000 times in just over two years.


While models are becoming cheaper, computing power is becoming scarcer. In March this year, Tencent Cloud, Aliyun, and Baidu Intelligent Cloud successively raised their computing power prices by around 30% within ten days. In the first half of the year, as AI agents boomed, the leasing price of inference computing power rose by over 40%.


The reason is that the way people use AI has changed. In the past, asking and answering a question only cost a small amount of tokens. Now, for an Agent to complete a task thoroughly, they need to bring context, adjust several tools, double-check back and forth, and start over if they make a mistake. Although the price per unit has decreased, the consumption has multiplied.


The price drop did not shrink the market; it brought in all the previously unquantifiable demand.



In 1981, IBM opened up the PC architecture, and compatible machines quickly flooded the market, driving hardware prices down. In the end, those who truly made a fortune were not the hardware manufacturers but the software companies like Lotus 1-2-3, WordPerfect, and later Microsoft.


Today, this path is being followed almost identically. While many are focusing on the slightly lacking underlying capabilities, companies above are already rushing to acquire users, entry points, and workflows.


American companies can study DeepSeek's architecture, use data generated by Kimi, or spend $20 million to train a decent open model. However, it's not easy to directly apply these.


Applications do not have downloadable weights or benchmarks. What determines whether they can run well depends on the tricks of a company accumulated over a decade.


In the past, major companies used to treat models as standalone products, waiting for users to open a specific chatbox. Now, models are starting to integrate into DingTalk, Feishu, and WeChat Work, with Agents directly embedded in the existing organization and workflow.


In the U.S., an office Agent might grow through Gmail, Slack, and Salesforce. In China, it might first integrate into DingTalk, Feishu, and WeChat Work, before connecting to financial software, supply chains, factory systems, and governmental platforms. Even if both sides use the same model, the final products will be different.


After the Price Cut


From 2022 to 2025, the U.S. kept tightening restrictions on chips to China. No H100, no B300, and even advanced processes and manufacturing equipment were restricted. Under these conditions, Chinese companies created the world's largest open model.


By 2026, the U.S. started discussing limiting the weight of Chinese models. OpenAI and Anthropic have reportedly expressed concerns to regulatory agencies, and Congress is also studying whether model weights can continue to flow across borders.


The blade has not fallen, but cries of pain are already heard among allies.


The chip ban targets the supply chain of a Chinese company, but the initial impact of the model ban hits the costs of American startups. The joint letter made it very clear that banning cannot stop the spread; it will only force developers to rush to download all available models before the door closes.


A chip is a physical entity that can block orders, logistics, and manufacturing. Once a model's weights are leaked, they can be endlessly downloaded, mirrored, quantized, fine-tuned, and distilled.


What's even harder to control is the price.


The United States still leads in the most cutting-edge proprietary models. OpenAI and Anthropic are still at the forefront, and even training the most top-tier models still heavily relies on NVIDIA GPUs, a fact that won't change in the short term.


However, when it comes to application, the outcome is rarely determined by a single state-of-the-art model. The competition lies in whether the model is affordable enough, if companies already have a ready-made digital entry point, and if developers can truly integrate into the business processes.


Looking back over the past thirty years, China has had many victories, but the winning formula has been largely the same.


Mobile payments, e-commerce, food delivery, live streaming—all these have involved transplanting externally developed technologies into the daily lives of over a billion people, turning them into tremendously scalable businesses. The foundational layer has long been provided by others. Transistors come from Bell Labs, x86 belongs to Intel, TCP/IP was funded by the U.S. government, and Windows, Android, and iOS all grew up on the West Coast. Even the most widely used frameworks in the era of deep learning, TensorFlow and PyTorch, originate from Google and Meta, respectively.


"Invented in America, scaled in China"—this statement has been made for three decades, with few true exceptions.


Now, an exception has emerged.


The technical documentation from the Thinking Machines Lab is very clear. One of Silicon Valley's most prominent and wealthiest new companies referenced the architecture of a Chinese company at the model's base layer and then trained using data generated by another Chinese model.


In the past, Silicon Valley has sourced production capacity, recruited talent, or sought markets from China. Now, it is starting to directly follow the technology path validated by Chinese companies.


The Silicon Valley admiration certainly has a history. For the past few decades, the most crucial underlying technologies of the computer industry have indeed mostly emerged from there. From the chips inside the machines to the IDEs used daily by programmers, Silicon Valley has long stood at the upstream end of the tech chain.


But in the end, the tech world looks at results. Who can propose a new architecture, who can bring down costs, who can make peers start to imitate, is eligible to redefine the upstream.


Now, the names of Chinese companies are starting to appear on this list.


Silicon Valley is likely to catch up in the future, as there are enough engineers, capital, and chips there. But applications that are already integrated into real business operations will not stop and wait for anyone. Whoever dives into the enterprise's processes first will have the opportunity to acquire customers, data, and the next round of improvements. The starting line has never been the same.


The era of AI applications will not choose an auspicious day to open with great fanfare. It first appeared in financial statements, in the skyrocketing number of API calls, in a programmer once again depleting their quota in the middle of the night.


Mass adoption is only a matter of time. This statement was made by an American.


Original Article Link


Recommended

Dalio Latest Interview: Already in an AI Bubble, 1% of Portfolio Is Bitcoin

Aug 4, 12:23
Dalio Latest Interview: Already in an AI Bubble, 1% of Portfolio Is Bitcoin

Palantir’s AI Story Starts to Show Up in Earnings and Contract Tables

Aug 4, 10:31
Palantir’s AI Story Starts to Show Up in Earnings and Contract Tables

Money Flowing Back to AI: Rewire News Morning Update

Aug 4, 10:14
Money Flowing Back to AI: Rewire News Morning Update

Robinhood Revenue Structure Undergoes Massive Shift: Predicts Crypto Trading Revenue to Surpass Stock Trading

Aug 3, 19:26
Robinhood Revenue Structure Undergoes Massive Shift: Predicts Crypto Trading Revenue to Surpass Stock Trading

12 Billion Shares Unlocked, Can SpaceX's First Quarterly Report Save the Stock Price?

Aug 3, 17:48
12 Billion Shares Unlocked, Can SpaceX's First Quarterly Report Save the Stock Price?

Morgan Stanley Analysis: Enterprise SSD Capacity Doubles, Storage Supply Faces Test

Aug 3, 16:40
Morgan Stanley Analysis: Enterprise SSD Capacity Doubles, Storage Supply Faces Test