Token Kills AI Applications

By: rootdata|2026/07/30 23:58:00

Author: Li Nan, Huang Xiaoyi, Yoky, Sun Rui, Silicon Star People

Editor | Wang Zhaoyang

By mid-2026, the cash on the books of the star AI video generation startup OiiOii still had nine digits in RMB, but it still initiated a new round of financing.

Previously, it signed an annual framework agreement with ByteDance's cloud service platform Volcano Engine: OiiOii needs to pay 50 million RMB to Volcano Engine by the end of 2026.

This made OiiOii's cash on hand a bit tight. The 50 million RMB would be used to purchase the "pureblood version" API of Seedance 2.0, ByteDance's most successful AI product—a multimodal video generation model.

The so-called "pureblood version" refers to the Seedance API provided through the official channel of Volcano Engine, which comes with full permissions.

After Seedance 2.0 amazed the AI industry and the entire film production sector, the market was flooded with Seedance APIs from various sources. Some obscure intermediary service providers mixed Seedance 2.0 with its lower configurations—Seedance 2.0 mini and Seedance 2.0 fast—and sold them at relatively low prices. Users wanting to generate a cinematic-quality blockbuster ended up with nothing but scrap.

The "pureblood version" Seedance 2.0 API is the lifeline for AI video generation applications. Just as large language models depend on top-tier computing power, namely high-end NVIDIA GPUs, the API of top models and the tokens generated from calling the API fundamentally determine the capabilities of AI applications—often referred to as AI Agents.

Thus, tokens from top models have become scarce and expensive hard currency. Seedance 2.0 is undoubtedly the most powerful video model currently on this planet.

According to official information, the price for generating 720p video without video input using Seedance 2.0 is 46 RMB per million tokens, with the cost for one second of video being approximately 0.99 RMB. In comparison, the cost for generating 720p video using Ling3.0 is 0.6 RMB, and Vidu Q3-pro is about 0.68 RMB, making Seedance 2.0 significantly more expensive.

OiiOii, which excels in generating comic and short dramas, needs the powerful Seedance 2.0, but spending nearly 40% of the company's cash flow to purchase its tokens sounds a bit crazy.

For today's AI application startups, this has long been commonplace. Compared to large language model startups (including OpenAI and Anthropic), which typically allocate 70% to 80% of their financing directly to purchasing or leasing GPUs and paying for cloud computing power, this is nothing unusual.

What OiiOii is purchasing is not the token consumption based on real business needs, but the first rights to the Seedance 2.0 API.

Few people notice that the official launch date of OiiOii on April 2, 2026, coincided with the day Seedance 2.0 opened for public testing to enterprises. This seems like Volcano Engine announcing OiiOii as one of the first distributors of Seedance 2.0, rather than OiiOii launching its own video Agent that day.

This is not important; what matters more is being the first to offer the "pureblood version" of Seedance 2.0.

The "pureblood version" of Seedance 2.0 comes bundled with a complete set of rights, such as access to face databases, maximum concurrency, etc. It directly determines whether the startup's product can run smoothly.

If an AI video generation application targets the consumer end, users need real-time generation, which requires sufficiently high concurrency. When the model is just launched, computing power is tight; without concurrency support from Volcano Engine, users of the Agent would have to queue, time out, or even be unable to generate videos.

An AI video application practitioner revealed to us that a 10 million RMB annual framework corresponds to about 400 concurrent uses, while a 5 million RMB annual framework corresponds to about 200 concurrent uses. The price difference translates directly to user experience differences.

Source: Volcano Ark

Users can access a larger inventory of face resources rather than the increasingly creepy few "AI faces" that are cycled through, which also follows the pricing of the Seedance annual framework. The difference between buying and not buying is the difference between product life and death; it seems you have no choice.

"Gross Loss Rate"

As the capabilities of top models determine the upper limits of AI applications, and with high prices "choking" AI applications, charging subscription fees to users based on token prices as a benchmark has become a "sub-landlord" business—a math problem that can never be solved, a variable function on a model pricing table.

Even so, the first rights to Seedance 2.0 obtained by OiiOii are not exclusive. Many of OiiOii's peers, such as LibTV and Flova, have also obtained the "first release" qualification for Seedance 2.0, almost simultaneously announcing that their products have integrated Seedance 2.0, with similarly high prices.

Volcano Engine is so generous; it will not grant any company exclusive first release privileges for Seedance 2.0, not even for the highest annual framework. Moreover, Volcano Engine's price floor is very strict: there are almost no discounts, and the token consumption for contracts in the tens of millions is returned based on usage, capped at 10%, meaning a maximum of 10% of the contract amount is returned as computing vouchers.

On the other hand, Volcano strictly prohibits reselling to prevent low-priced tokens from entering the gray market.

Under this mechanism, a harsh reality emerges: AI video tools are almost impossible to make money from fully-fledged Seedance 2.0 tokens.

Theoretically, selling Seedance 2.0 tokens to customers at the original price, without considering rebates, yields a maximum gross profit margin of 10%. However, in reality, no AI video generation tool sells Seedance 2.0 tokens at a higher price—even at the original price. After all, theoretically, anyone can do business directly with Volcano Engine. If intermediaries can profit, why not negotiate directly?

According to reports from "Intelligent Emergence": comic and film production companies also mostly purchase "fully-fledged" Seedance 2.0 APIs directly, including many leading companies in the industry, with the highest single recharge reaching 50 million. This fully demonstrates the market's thirst for Seedance 2.0 and implies that AI application startups with the highest rights to Seedance 2.0, if they do not offer competitive token prices externally, will find no users.

"Products with tool attributes do not have loyalty," a VC person told us.

Thus, a "wild path" emerged: leading AI application companies began to actively lower prices to acquire customers, even starting to resell tokens directly.

LibLib's parent company, Evoken, is the largest AI application company in China in terms of token procurement volume. It started from the open-source model community and quickly transitioned to a multimodal creation platform. At the same time, it aggressively conducted multiple rounds of financing, recently completing a $300 million Series B round, with a post-investment valuation exceeding $2 billion, forming a comprehensive barrier of capital and product matrix—LibLib AI is a creation community and image generation platform, while Lovart and Xingliu target overseas and domestic design markets, respectively, and LibTV focuses on professional video production...

However, Evoken has not developed any models itself, not even retrained models. This has led to frequent debates about whether this company is merely an intermediary for top models without any technical barriers. But undoubtedly, Evoken is a major customer of top model tokens.

In terms of video tools, LibTV is a top-tier customer of Seedance 2.0, signing a large annual framework of 50 million for use as ammunition for LibTV Agent. In design tools, last year, Nano Banana exploded in popularity, and Lovart, under LibLib, also signed a multi-million dollar contract with Google Cloud.

However, according to estimates from third-party platforms like Similarweb: in June 2026, Lovart's website had about 3.1 million monthly visits, while LibTV's website had about 2.2 million monthly visits. This level of traffic is already considerable, but visits do not equate to creation. Public information shows that LibTV has several hundred thousand paying users, which, although exceeding most peers, basically confirms that it cannot consume all the Seedance 2.0 tokens purchased at great expense by the company.

Volcano Engine does not accept "returns" of tokens, which has led to the emergence of side businesses for AI application companies.

Evoken's founder and CEO, Chen Mian, mentioned in an interview with "Latepost" that there were rumors that LibTV was reselling Seedance 2.0 API Keys for profit, which he explicitly denied.

What we learned is that LibTV's resale of Seedance 2.0 API Keys indeed does not "make a profit" because it is sold at a loss.

Some of LibTV's competitors, engaged in AI video generation, told us that LibLib has repeatedly attempted to sell Seedance 2.0 APIs to them at prices cheaper than going directly to Volcano Engine: directly purchasing from Volcano, contracts over 10 million RMB can receive a 10% computing voucher return; but LibLib can also provide them with an additional 10% return on top of that.

There is certainly no price difference, let alone profit, but it generates revenue. It can boost LibTV's ARR (Annual Recurring Revenue), providing numerical support for further financing narratives.

This type of ARR is essentially "negative gross profit" revenue: obtaining tokens at a 10% discount contract price, then selling them at a further 10% discount to friends needing lower-priced tokens, structurally represents a loss-making transaction.

What it generates is not even a gross profit margin, but a "gross loss rate."

On the other hand, LibTV and Lovart's pricing strategy for paying users also makes its gross profit margin as an AI application company seem illusory.

Let's do a rough calculation...

The average monthly price for LibTV's standard annual membership is 47.4 yuan, with the official estimate generating a 56-second Seedance 2.0 720P video, equivalent to approximately 0.85 yuan in member revenue per second. Assuming users utilize the entire quota for Seedance 2.0 without video input, and estimating based on the Volcano Engine's price of about 0.99 yuan/second, the marginal gross profit margin is approximately -17% when only considering model invocation costs. Even if 10% of the computing vouchers can be fully redeemed, the marginal gross profit margin would still be around -5%. This also needs to cover night discounts, card opening bonuses, and limited-time free offers.

Lovart's situation is similar. If all points are used to generate medium-quality 1K images with GPT Image 2, referencing the "most popular" Pro package annual benefits, the current "end-of-month carnival" discount price of $39, and OpenAI's official API standard pricing, when users fully utilize the quota, the corresponding marginal gross profit margin is about -87% (a few days ago, the discounted average monthly price for this package was still $45, corresponding to a marginal gross profit margin of about -60%).

However, these are all part of a "stress test." In reality, users may not fully utilize their quotas and may also use points for lower-cost models and generation schemes. The platform may also obtain varying degrees of discounts when purchasing different models' APIs, which could narrow the actual losses.

Chen Mian, CEO of Yanyu Technology, also responded to this point in an interview with "Latepost": If users utilize all tokens of the package quota, LibTV will definitely incur losses. But if a user purchases 1 million points and only uses 200,000, the remaining 800,000 becomes profit. There are also black market users who reverse-engineer LibTV, turning annual membership accounts into APIs, leading to losses, so governance is necessary.

The so-called "reverse engineering" has a clear example: In February 2026, the operator "Chengzhu" of the influential WeChat public account "Web 3 Sky City" in the Web 3 and AI fields publicly criticized Lovart, a visual design application under Yanyu Technology, claiming to have paid a $540 annual fee for Lovart Pro membership, but the account was banned 10 days later without a refund. "Chengzhu" appealed via email and proactively communicated with Lovart team members but received no response.

According to our understanding, "Chengzhu"'s membership account was deemed by Lovart's internal risk control mechanism to have engaged in "reverse engineering," and the company has upgraded its governance measures. The decision to refuse further communication with "Chengzhu" and deny a refund was made by Chen Mian himself.

This incident sparked considerable controversy, leading to discussions about the prepayment mechanism of AI application software, the transparency of account banning standards, and consumer rights protection. It highlights how fragile a business model relying on users paying without consuming token quotas can be; moreover, in an environment where token prices are extremely sensitive, there are many loopholes that can be exploited for reverse arbitrage.

In fact, the true reverse engineering, packaging unused membership tokens into APIs for sale, is being done by Lovart or LibTV themselves.

Because unused tokens by users do not truly convert into profits, AI application companies have already paid fees to model vendors for these quotas.

They are included in the annual framework. The logic of strong model vendors is: you can buy more when you run out, but you absolutely do not get refunds for unused tokens. The so-called 1 million points used 200,000, with the remaining 800,000 turning into profit does not exist. In other words, the "unused" token quota actually becomes the "inventory" of AI application companies.

Since it is inventory, it must be cleared out; otherwise, losses will increase. This explains why other AI video generation teams receive offers from LibTV to sell tokens at discounted prices. The official crackdown on users engaging in "reverse engineering" to resell tokens is actually a competition for the resale rights of tokens with users.

Compared to the losses caused by selling off token "inventory" at discounted prices, the losses from users fully utilizing their rights are greater. Even if the notion that "unused tokens are profit" holds, it locks in the upper limit of the profit model—when all users recharge but do not consume a single token, theoretically, its profit can be maximized.

When a user consumes one more token, it moves closer to losses. This creates a dangerous adverse selection: the lower the price, the easier it is to attract heavy users who consume expensive models, leading to higher point usage rates and leaving less residual income for the platform.

This cycle may expand the losses of AI application companies, with the only benefit being a temporarily better-looking ARR.

AI application companies are well aware of this, but when presenting to VCs, they can still hide these "traps." Some VCs have revealed to us: the gross profit margin that LibLib presents to investors and the capital market is likely positive.

This tests "financial skills": bonuses, limited-time free offers, night discounts... all these rigid costs that truly eat into gross profits can be counted as sales rates, and the early high sales rates can be euphemistically called "growth investment"; all income—including user purchases of membership points and income from reselling tokens—can be counted at original prices, making it not impossible to achieve "positive gross profit margins" under different metrics.

In our discussions with many investors and entrepreneurs, a consensus is that even when calculated at the broadest scope, the gross profit margin range of AI applications centered on consuming and selling tokens is generally only 10%-25%, which is almost insufficient to cover various costs.

Most Chinese AI application startups have been expanding into overseas markets since day one, but informed sources have calculated that the cost of acquiring a paid user overseas is about $200, while users pay an average of $50, and the token gross profit cannot even cover user growth, let alone the expenses for production research, operations, and management.

In recent years, the capital market and startup circles have preferred to understand AI applications through the SaaS framework. But now investors have awakened to the fact that "SaaS means the more users, the cheaper it gets, while AI applications mean the more users, the more you lose."

This points out the most essential difference between AI applications and traditional software—the marginal cost of traditional SaaS dilutes with scale, while the marginal cost of AI applications re-emerges with usage. The more users love to use it, the more tokens are consumed; the more tokens consumed, the heavier the costs.

At the core, AI application companies operate a "continuously loss-making token turnover chain," which is the "gross loss rate."

"Token Controllers"

AI application startups are not knowingly jumping into a pit of fire with tokens; they are also caught in chaos.

The apparent cost of purchasing tokens from model vendors is clear, but product pricing, user subsidies, and gross profit calculations are all driven by the post-settlement system from the models. The initially thought gross profit turns into a gross loss rate when the bill comes out.

Taking Seedance 2.0 as an example: the actual cost of generating a video at once will simultaneously consider the input video length, output video length, output resolution, frame rate, model version, minimum billing duration, and duration tiers. The API for video models may charge uniformly for 1 to 4 seconds at 4 seconds, and the growth between different lengths is often not linear. After signing a contract, AI application companies incur consumption daily, but usually, it takes T+N days to see the detailed costs in the bill.

This creates a typical state of being "trapped in tokens": the company knows the apparent costs of tokens but does not know how the real costs will be magnified in what way, at what time, and to what extent.

However, AI application startup teams cannot escape this dependency; instead, they are sinking deeper. After all, they cannot do without the "benefits" brought by top models.

If we present the cost sources of AI applications in a formula, it would be: "Real generation cost = Single generation price × Average draw frequency." This applies to videos, images, and PPT designs alike.

Top models have a high generation success rate, reducing the draw frequency, which can sometimes make the underlying calculation costs lower.

How should AI application teams choose between high upfront costs, good effects, fewer draws, and not necessarily the highest overall costs of top models, versus lower upfront thresholds, cheaper calls, more draws, and increased overall costs that may sacrifice conversion retention of suboptimal models?

Industry insiders have revealed: to increase orders, clients of Keling 1.0 could receive discounts as low as 3.5%. Such a discount is already a loss for Keling's inference costs, and with later upgrades of the Keling model, inference computing costs rise, making this price a double loss. Its only purpose is to seize customers.

The objective result is that those model vendors trying to sell more tokens to application companies at lower prices will still be sidelined. A side example that can confirm this is that Seedance 2.0's daily request volume on OpenRouter remains stable at over 3,000, while Keling 3.0 only has 300-400 requests.

Therefore, to solve the inherent "gross loss rate," there are two direct methods: raise product pricing or lower token costs. Currently, most AI application startups are choosing the latter.

On the surface, everyone is actively iterating products, updating features, and creating differences, but in reality, a significant portion of their energy is spent looking for cheaper tokens. And cheap tokens come from various sources.

First, the low-price dumping by "upstream" AI Agent companies is an important channel.

As mentioned earlier, LibLib sells tokens at an 8.1% discount to smaller AI application teams that cannot afford the full version of Seedance 2.0. Some startups welcome such "discounts," while others hesitate to take them for fear of violating Volcano Engine's sales policies.

Secondly, peers undergoing business transformations can also provide low-cost tokens.

As previously mentioned, leading model vendors have strong annual framework policies with guaranteed quotas, and "expire without waiting"—if not consumed within a year, the remaining token quota is cleared. This is also why reselling tokens is repeatedly banned by model vendors; whether for loss reduction, cost recovery, or "creating" gross profits, it determines that AI application teams must do so.

Especially when some companies burn money and fail to tell stories, becoming unable to continue or on the verge of bankruptcy, their remaining "huge" token quotas may be sold off at extremely low prices.

An interesting detail is that some companies' token procurement costs exceed tens of millions, but the main purpose is not to conduct business but to tell a story: to inform the capital market that they have become major clients of top model vendors, showcasing their strength and quickly raising another round of funding. Some VC insiders have told us that buying tokens has completely "detached from reality" for some startup teams.

Additionally, seeking government subsidies is a common and more reliable channel for obtaining lower-priced tokens.

Local governments at the district level will provide computing power or model subsidies to AI companies recognized in their jurisdictions, allowing application companies to reduce token costs to 75% to 90% of the original price. However, such resources are not widespread, and to obtain cheaper tokens, everyone must "cross the sea by relying on their own abilities."

Ultimately, everyone is trying to find ways to obtain cheaper tokens. This maximization of the pursuit of cheap tokens will inevitably lead to the proliferation of "intermediaries."

As an underground AI industry, intermediaries essentially wholesale tokens for retail. They can be roughly divided into three categories.

The first category is "legitimate" resale, which involves large-scale purchases of overseas model tokens, obtaining discounts, and then reselling them domestically. This is still considered a "normal" distributor logic.

The second category is pure black market operations, often run by individuals, with offers like "one million tokens for three dollars," which can disappear at any time or be shut down.

The third category is "pooling" resources. These intermediaries operate by collecting token quotas from various scattered channels and creating a unified routing layer.

Intermediaries are one of the main channels for low-cost tokens in the AI application market. However, the prices of these resources vary greatly, and their quality is unstable. In video generation, the phenomenon of "watering down" is commonplace.

One intermediary supplying a certain AI startup mentioned to us: "For example, when calling Keling's API, 8 out of 10 times it is Keling, and 2 times it is this company's own model. If the intermediary wants to make more profit, they adjust this ratio."

AI application companies are well aware of this, but some still "tolerate" watered-down tokens in their products, further reducing the service quality received by end users. After all, "failed card draws" are the norm, and even if users suspect that what they are using is not up to standard, they can provide an explanation.

Even with these various channels, AI application companies are still looking for ways to further reduce costs.

To legitimately lower token costs and find some gross profit space for financial statements, some startups have invented another path: designing their products as precise token controllers—by limiting duration, tiering packages, degrading models, hiding parameters, and other methods, they weaken the impact of tokens on gross profit.

Specifically, some applications directly limit the duration of videos generated by users, such as uniformly generating 4-second videos to avoid wasting resources by charging for 4 seconds when only 3 seconds are produced.

The "package trap" is a more clever approach; some applications do not provide the best models in mid-tier packages and do not open the optimal models for high-consumption scenarios. They also design packages that users "cannot fully utilize," hoping for user inactivity and low usage to gain arbitrage space.

They use formulas like "199 yuan tier user count × 2 billion token reserve" to negotiate wholesale prices with model vendors, obtaining procurement prices far below actual consumption needs. The unused portion is then sold as an intermediary.

A VC complained to us: during their due diligence on a company, they discovered these "inflated figures" and subsequently asked whether their users could actually consume that many tokens, while the founder only emphasized the number of users purchasing that tier of package, avoiding discussion of actual consumption.

Investors were not pleased; a business where user activity and gross profit margins are inversely proportional is not a good business.

AI applications "exist for tokens," which raises a fundamental question that shakes the foundation of the company: products are no longer designed solely around user experience but increasingly around "how to make users consume fewer tokens."

Of course, there are also engineering cost-reduction methods that consider user experience. For example, generating low-resolution versions of videos using cheaper models first, then enhancing clarity through super-resolution technology. Some AI tools reduce the number of card draws through prompt words, workflows, parameter control, post-processing, caching, and model routing, thereby saving tokens.

Among them, OiiOii is one of the companies that has designed its "token controller" quite cleverly.

OiiOii refers to the points purchased by users as "boxed meals." Unlike LibTV, which aggressively competed for customers at 30-40% lower prices, OiiOii's main model (with multiple images for reference, 10 seconds for 140 boxed meals) calculates a price of 0.96 yuan/second, plus high-definition export, bringing the overall price to about 0.99 yuan/second, roughly equivalent to selling Seedance 2.0 tokens at a fair price, with a barely squeezable gross profit space, although still thin.

It can do this due to some clever "mechanism" designs:

The interface does not display video specifications and point deductions but directly pushes users into the script—character—scene—shot—video creation process, neither giving users the chance to feel "pain" at any time nor informing them whether they are using Seedance, Keling, Sora 2, or even other cheaper models during the early creation process. Only at the shot step do users get to choose the model. Users cannot optimize consumption or compare prices with other companies; they can only temporarily forget the price. "Opacity" is its gross profit protection mechanism.

However, OiiOii dares to set such pricing and design such token mechanisms because it has a good reputation and recognition among specific professional video production groups like comic dramas. Beyond the life-and-death game of "colluding with tokens," it has established a certain uniqueness.

A BD employee from a cloud service provider that interacts with short drama and comic drama clients frequently told us: OiiOii has a certain reputation in the short drama and comic drama circles, and recently, LibTV has also caught up. These clients mention that LibTV's workflow is user-friendly and the canvas is easy to use.

Previously, the core reason for LibTV's attention was its controversial "cost-performance ratio" and "gross loss rate." However, with product improvements and breakthroughs, LibTV, which has always been marked by token price reductions, has also gained some confidence to raise prices.

In June, LibTV Agent was launched. It advanced AI video from generating single shots to delivering complete films, automatically completing planning, shot composition, generation, and editing through professional skills while retaining storyboards and node workflows for manual modification; the 3D director's platform, character library, and long-term memory further improved the consistency between characters and shots.

Then, LibTV made a rare and bold price increase.

Previously, the annual subscription prices for the two tiers of LibTV's supreme version were 8499 yuan and 10999 yuan, respectively, and now both have increased by 1000 yuan. The corresponding calling price for Seedance 2.0, which was previously as low as 0.37 yuan per second, is now at least 0.41 yuan per second.

This has led to complaints from some price-sensitive heavy users, but with improvements in product experience, this is also an inevitable measure.

In addition to raising prices, there is a more direct approach: directly targeting overseas markets with higher prices. To improve gross profit margins, almost all Chinese AI applications prioritize serving the global market, especially regions like Europe, America, and Japan, where users have a higher willingness to pay. Lovart, under Yanyu Technology, is also an active experimenter.

According to previous reports from Elsewhere: Lovart was deeply inspired by Manus. When Manus's model capabilities crossed a critical point, it quickly packaged them into a product with a good experience, while Lovart learned that GPT-image-1 would be released in three weeks and pulled the best people from the entire company to Shanghai for closed development. Manus completed a cold start by actively contacting Silicon Valley tech influencers, while Lovart leveraged overseas tech bloggers for dissemination. A user-generated poster of the Tesla Cybertruck even received a like from Musk.

However, Lovart's overseas expansion has not been smooth sailing. As early as May 2025, when it launched, Lovart established a local team in Silicon Valley responsible for product user growth and market operations, but this team was filled with turmoil. By the end of 2025, American team members, including Lovart co-founder and COO Elena Leung and co-founder Aimee Yang, had all left, and the current team members are all newcomers who joined in the first half of 2026.

In the pursuit of higher "net value" users in the global market, improving product engineering and design to raise prices, and finely setting the "token controller" of AI applications, more and more AI application startup teams are making difficult trade-offs. Designing a finely tuned token controller is indeed the most relatively controllable option in terms of results.

However, if all AI application startups must first design themselves as token controllers, they will never be able to provide a true user experience or genuinely enhance product strength. In the larger market competition, this may make them more passive.

The most obvious threat comes from the giants.

When powerful internet giants with strong product capabilities leverage their model advancements and the capabilities of third-party models, using their financial strength and fearlessness of token consumption to quickly establish their own AI application product matrix, the prospects for startups are squeezed. The rise of Tencent's WorkBuddy and the overflow of product capabilities behind it is a prime example that AI application entrepreneurs should pay attention to and be wary of.

Since its public beta in March 2026, Tencent has expanded exposure through WeChat information streams, public accounts, subway and office building advertisements, as well as measures to lower user trial costs such as registration bonuses, sign-in rewards, model exemptions, and double credits, attracting a large number of users for WorkBuddy. The simultaneous high-frequency iterations further solidified WorkBuddy's position in the AI productivity wave.

In June 2026 alone, this product completed at least 17 version releases, sometimes even two updates in one day. In addition to quickly fixing bugs, WorkBuddy also rapidly integrated with Tencent Docs, WeChat Work, and other Tencent ecosystem applications, enhancing capabilities such as multi-agent collaboration, cloud tasks, and enterprise project management, ultimately becoming an AI star within Tencent that surpassed Yuanbao.

When releasing the first quarter results in May, Tencent bluntly stated that in terms of daily active accounts, WorkBuddy has become the most popular efficiency AI service in China. According to data from Analysys, from March to June, WorkBuddy's monthly visits on the PC side grew from 8.85 million to 20.97 million, more than doubling.

This is an achievement that any AI application company would envy but cannot attain. WorkBuddy can tap into the entire Tencent traffic resource, and more importantly, it does not have to manage the switches of token controllers as meticulously as startups do.

Finding a Way Out

More and more AI application entrepreneurs are beginning to explore alternatives to tokens.

Star entrepreneur Zhang Yueguang posted on Xiaohongshu in June: To evaluate whether an AI application is viable, the real top three indicators are paid user gross profit margin, renewal rate, and CAC (customer acquisition cost). To achieve a healthy balance of the first two, either the business model of selling tokens must be transformed into another business model, or tokens must be sold to users who are not sensitive to token prices. "All products that sell tokens based on sensitive users can only be the business of model suppliers."

Zhang Yueguang's own entrepreneurial product corresponds to the two different solutions he is looking for.

One product is Xingmian, which attempts to cover token costs through game and IP revenue; the other is Dokie, which aims to have users pay for usable results rather than comparing the unit price of each model call. Additionally, according to his public updates, Dokie has completely reduced the free trial quotas for over 140 countries to zero, which has virtually no impact on paid users.

For a long time, "token Maxxing" was regarded as a standard, with the industry fervently pursuing token consumption, viewing it as a symbol of product strength and a positive market outlook. However, people have finally realized that the pursuit of token consumption is misguided; the gross profit structure is the true guiding principle. In other words, companies that truly thrive on tokens are not those that burn more tokens but those that can sell tokens in a different way.

The most typical example is in the "hardware + application" sector, with Plaud being the most representative.

According to our understanding, Plaud's non-hardware AI service net profit has reached an annual scale of tens of millions of dollars. Its total sales target for 2026 is $500 million, with over 2 million installed users.

This is due to a different payment structure.

First, there is the hardware itself. In the consumer electronics industry, the retail price of terminals usually needs to reach about three times the BOM cost, which is the material cost, for companies to have sustainable operating space. The Plaud Note Pro is priced at approximately $179-189, with a BOM cost of about $60, resulting in a hardware gross margin of about 67%, which aligns with this principle.

More importantly, there is AI subscription revenue. Its AI subscription is divided into two tiers: Pro annual payment is approximately $8.33 per month, and Unlimited annual payment is approximately $20 per month. This pricing is neither too expensive nor too cheap; the key is that its AI scenarios are light, mainly involving transcription and summarization, not video generation or agent reasoning. According to an estimate from a Plaud supplier: based on subscription revenue of $8-20 per month per person, Plaud's AI processing cost is about $1-2 per month per person, allowing Plaud's software gross margin to reach approximately 75%-90%.

Overall, the situation for hardware is relatively better than for software, as there is no direct competition with model companies or model products, and there is room for premium pricing.

An AI hardware entrepreneur has categorized market products into several types:

  1. Hardware self-sufficient type. By pricing three times higher domestically and five times higher internationally, hardware can avoid losses, and subscriptions can provide additional profits.

  2. Low-margin hardware + subscription supplement type. Plaud falls into this category.

  3. Financing growth type, which many early AI hardware startups belong to; finally, there is the low hardware price + no subscription type. For example, most domestic AI companionship or AI toy products mainly earn a bit from hardware to subsidize token costs.

However, from this entrepreneur's perspective, hardware is also becoming competitive, and pricing above three times the BOM cost is becoming unsustainable. "Gross margins are too low; later on, they will be eaten away by channels, after-sales service, and returns; if we only pursue hardware gross margins, we might raise the initial purchase price too high, affecting shipment scale and subsequent subscription conversion." If this is indeed the case, then AI hardware must also seek alternative paths.

Beyond the problem-solving methods of hardware companies, some more general ideas are emerging.

A recent trend is that more and more AI applications, including hardware companies, are beginning to "post-train" their own models.

The founder of the recently popular hardware product, "AI Mushroom" Swoii, known as "Uncle He," revealed to us that they are using self-developed models to achieve low-cost personalization, meeting user UGC scenario needs and improving user satisfaction and willingness to pay.

Mindverse, an AI application company that gained popularity with Macaron, is transitioning to a model research company. Its Mind Lab has just released and opened the Macaron-V1 model weights, serving its own products while also beginning to provide model and post-training capabilities to more B-end applications.

Additionally, we understand that well-known serial entrepreneur, Teambition founder, and former Vice President of Product at Feishu and responsible for Doubao's PC product, Qi Junyuan, is starting a new venture. His team is developing and internally testing an AI Agent product, which is largely based on self-developed edge models to minimize reliance on model vendors' APIs and tokens.

Although these products have different scenarios, there is an important common logic behind the differences: more and more entrepreneurs dedicated to creating good AI products realize that post-training a model and integrating it with applications is much more cost-effective than simply purchasing tokens from model vendors.

Especially those AI application companies with clear scenarios, rich data, and stable calling scales have the opportunity to replace expensive general SOTA model APIs with specialized models through post-training. As the calling scale expands, the one-time investment in model training can be amortized over the long term, significantly reducing the unit token cost.

If this path can be successfully navigated, AI applications will become more popular. As for now, pure software AI projects have already cooled down in the capital circle.

Some investors have indicated that in the first half of this year, the focus of some first-tier funds has shifted to embodied intelligence and hardware. This does not mean they are not investing in software. They are still looking at the To B direction, especially specific scenario applications targeting "wealthy small bosses" are favored. a16z continuously invested in several AI projects related to dentistry, medical operations, and home services in June. In the spring project of Qiji Chuangtan, a batch of similar companies targeting factories, brand owners, and overseas enterprises also emerged.

This shift is not difficult to understand. Although "wealthy small bosses" care about price, they care more about whether tools can help them produce more and close more deals. As long as the benefits created by AI significantly exceed the usage costs, even an annual fee of tens of thousands or more is worthwhile. The AI businesses serving them may be "small," but they can genuinely make money in business.

Moreover, the rising trend of FDE is also a reflection and breakthrough of the industry's token dilemma in AI applications. FDE shifts the product value anchor from "how many tokens were called" to "what was accomplished for the customer," giving startups a chance to charge based on delivery results. When tokens no longer determine the upper limit of product value, AI applications have the opportunity to gain more profits.

Ultimately, tokens only buy an entry ticket, not a moat. If one cannot escape, they can only be devoured by tokens.

This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.

You may also like

iconiconiconiconiconiconicon
Customer Support:@weikecs
Business Cooperation:@weikecs
Quant Trading & MM:bd@weex.com
VIP Program:support@weex.com