option
Home
News
U.S. AI Models Mirror DeepSeek: Worse Performance, Higher Costs, Yet Tailored for American Compliance

U.S. AI Models Mirror DeepSeek: Worse Performance, Higher Costs, Yet Tailored for American Compliance

September 3, 2026
42

U.S. AI Models Mirror DeepSeek: Worse Performance, Higher Costs, Yet Tailored for American Compliance

Former OpenAI Chief Technology Officer Mustafa Suleyman’s inaugural project has unexpectedly flipped the script on large language model development between China and the United States. On July 15, he launched Thinking Machines Lab and unveiled its first model, Inkling. Leveraging a hybrid expert architecture inspired by DeepSeek-V3 and utilizing synthetic data from open models like Moonshot’s Kimi K2.5 for post-training initialization, the results are nuanced. Across multiple benchmarks, Inkling trails behind Kimi and GLM, with pricing that is notably higher. Despite raising $2 billion and achieving a $12 billion valuation, the debut fell short of hype, yet Suleyman may never have aimed for the top tier from the start.

Examining the hardware, Inkling matches the prestige of a star lab. It employs a hybrid expert architecture with 975 billion total parameters, activating 41 billion per token. Pre-training utilized 45 trillion tokens across text, images, audio, and video, supporting up to 1 million tokens of context. The model processes text, images, and audio, outputting text while allowing adjustable reasoning intensity to balance performance, speed, and cost. With weights released under the permissive Apache 2.0 license, developers can self-deploy or fine-tune via Thinking Machines’ Tinker platform. The strategy is clear: rather than targeting consumers directly, Inkling serves as a foundation for enterprise AI applications.

The technical report reveals a critical detail in its architecture. Thinking Machines explicitly states that Inkling’s hybrid expert design follows DeepSeek-V3, featuring numerous expert modules where each token activates only a subset, along with DeepSeek-V3’s no-assistant loss load balancing. Chinese influence extends to training: post-training began with supervised fine-tuning using synthetic data from open-weight models, specifically citing Kimi K2.5.

A clear roadmap emerges: an American lab founded by OpenAI’s former CTO, its first model mirrors DeepSeek’s architecture, and post-training relies on Kimi’s data. Borrowing from open models is not inherently problematic, as DeepSeek and Kimi publicly share weights and usage rights, building upon public achievements. More strikingly, despite absorbing Chinese model technologies, Inkling fails to surpass its predecessors. Thinking Machines candidly admits: Inkling is not currently the strongest model, whether open or closed-source.

The numbers highlight the disparity. In final human exam pure text evaluations, Inkling scored 29.7%, while Kimi K2.6 and GLM 5.2 achieved 35.9% and 40.1%, respectively. With tools enabled, Inkling rose to 46%, still lagging behind Kimi K2.6’s 54% and GLM 5.2’s 54.7%. In SWE-Bench Pro, testing real software engineering capabilities, Inkling scored 54.3%, compared to Kimi K2.6’s 58.6% and GLM 5.2’s 62.1%. The gap was most pronounced in Terminal Bench 2.1, where Inkling scored 63.8%, while Kimi K2.6 and GLM 5.2 reached 71.3% and 82.7%, respectively.

Inkling excels in web application design, audio understanding, and security evaluation, even outperforming Kimi and DeepSeek in certain math tasks. However, it has not yet reached the top tier for comprehensive reasoning, programming, and agent capabilities. Pricing offers no relief: on the Tinker platform, the 64K version charges $1.87 per million Prefill Tokens and $4.68 per million Sample Tokens, already discounted by five days. The 256K version costs $3.74 and $9.36, respectively. Compared to Kimi K2.6’s API rates of $0.95 input and $4 output, and GLM 5.2’s $1.4 and $4.4, Inkling’s prefill price is nearly double Kimi K2.6’s, with higher generation costs and no obvious price advantage. The situation is delicate: the architecture references DeepSeek, post-training relies on Kimi, performance lags behind China’s leading open models, and pricing is higher.

The intrigue lies in Thinking Machines’ unique background. The company resembles an OpenAI alumni network—introduced in February 2025, the team of about 30 people included two-thirds from OpenAI, with others from Meta and Mistral. Founding members included John Schulman, Barret Zoph, Lilian Weng, Andrew Tulloch, and Luke Metz, with Reuters reporting Suleyman recruited at least 20 researchers from his former employer.

Suleyman joined OpenAI in 2018, contributing to DALL-E, Codex, ChatGPT, and Sora, becoming CTO in 2022, and briefly serving as interim CEO after Altman’s dismissal in November 2023. This group, intimately familiar with OpenAI’s development and productization, left to build a new technical system, leading outsiders to expect an OpenAI-inspired route. Instead, the primary technical source points to Chinese models, signaling a dramatic shift: those who knew OpenAI best ultimately adopted the technical blueprint of Chinese models.

This conclusion requires careful analysis. OpenAI’s recent cutting-edge models are closed-source, preventing external access to weights, architecture, or training plans. Even if Suleyman knew internal technologies, he could not transfer OpenAI’s commercial secrets to the new company. Conversely, DeepSeek and Kimi have publicly released weights and technical reports, allowing their architectures, training methods, and toolchains to be legally researched and reused.

Inkling is an open-weight model, and for companies building an open ecosystem, adopting mature solutions from open models like DeepSeek and Kimi is a logical engineering choice. Developing an open model while ignoring publicly available Chinese achievements would incur higher trial-and-error costs. Thus, drawing inspiration from Chinese models does not prove Suleyman believes DeepSeek surpasses OpenAI. The differences in openness and reference conditions are distinct, but the signal is clear: Chinese models have become a key reference for American AI teams in the open-weight space.

The real question is why Thinking Machines, with existing architectures and training experiences, produced an Inkling with weaker performance and higher prices. Suleyman has a top team, ample funds, and the latest NVIDIA training systems. There is no reason to overlook these gaps, suggesting the high-cost alternative may be a calculated decision.

Considering Thinking Machines’ treatment, Inkling should be more than a decent model. In July 2025, the company completed a $2 billion seed round before launching any product, valuing it at $12 billion, with investors including a16z, NVIDIA, AMD, and Sequoia. Four months later, negotiations aimed for a valuation up to $500 billion, positioning the company as a potential competitor to OpenAI and Anthropic. Launching a product that underperforms Chinese models and costs more falls short of expectations. Yet, from a business perspective, Suleyman may never have intended to win benchmarks.

The company emphasizes open weights, multimodality, and customizability, encouraging enterprises to use Tinker to fine-tune Inkling into specialized models for customer service, coding, and industry agents. Suleyman bets on a market that does not prioritize benchmark rankings. Enterprises care more about private deployment, data control, and continuous training according to their business needs. Much of this demand is met by Chinese open models: OpenAI and Anthropic remain closed-source, and after Meta’s Llama4 underperformed, it reduced its open-source efforts. Meanwhile, DeepSeek, Kimi, Qwen, and GLM continuously release weights, with performance nearing American closed-source models at much lower prices.

US enterprises have quickly responded—according to OpenRouter data provided to CNBC, since February 2026, the weekly share of Chinese models in tokens called by US enterprises through the platform has exceeded 30%, reaching 46%, compared to an average of 4.5% in the first half of 2025. Cursor tested multiple base models for Composer2, ultimately choosing Kimi K2.5 for its strong evaluation. Bridgewater Fund fine-tuned Alibaba’s Qwen through Tinker, achieving better results than some top closed-source models. Chinese open models have become a cost-saving shortcut for US enterprises, but this path is now under pressure.

Simultaneously, the US regulatory atmosphere toward Chinese AI models is tightening. Relevant departments have issued warnings about data security, model sources, and supply chain risks, with some companies using Chinese models facing investigations. Even without unified restrictions, policy uncertainty affects enterprise choices, especially for those handling government business and sensitive data.

This creates a market space for Inkling: an open-weight model from an American company, using the Apache 2.0 license, supporting private deployment and customization. US enterprises choose it without worrying about political controversies from using Chinese models, making it easier to pass compliance reviews from government clients and large corporations. Inkling’s real value lies in this compliance certainty—Suleyman transformed public achievements from DeepSeek and Kimi into a model provided by an American company. It does not need to defeat Chinese models in benchmarks, only to be an acceptable choice for US enterprises hesitant to continue using Chinese models. This explains its higher pricing: policy has narrowed the competitive scope, making performance and price gaps less critical. For some US enterprises, a slightly weaker and more expensive model is sufficient.

Related article
Meta Removes AI Photo Editing Feature Following User Backlash Meta Removes AI Photo Editing Feature Following User Backlash Meta, the social media giant, is once again embroiled in a public debate regarding the delicate balance between artificial intelligence and user privacy. According to TechCrunch, Meta’s Superintelligence Labs introduced a new AI image generator, Muse
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m. DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m. Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage
Related Special Topic Recommendations
Video creation AI Short-Form Video Hook Generators for Reels, Shorts, and Product Launch Campaigns
AI Short-Form Video Hook Generators for Reels, Shorts, and Product Launch Campaigns

2026 Latest Best AI Short-Form Video Hook Generators for Reels Shorts and Product Launch Campaigns! XIX.AI curates top-rated powerful game-changing tools that deliver must-try results through real-world tests. This weekly updated page offers free vs paid comparison rankings to help you find the perfect solution for boosting content creativity and productivity. Explore now to Unlock your AI edge!

11 tools
xix.ai
writing AI Article Outline Tools for Long Form Writing
AI Article Outline Tools for Long Form Writing

2026 Latest Best Top-Rated AI Article Outline Tools for Long Form Writing! This curated collection features powerful, game-changing tools that deliver accurate, structured outlines with real-world tests to boost writing efficiency significantly. XIX.AI is part of this elite selection. Get a free vs paid comparison to help you choose the perfect tool. Explore now and Unlock your AI edge!

9 tools
xix.ai
chatbot Best AI Roleplay Chat Apps for Language Practice, Interview Prep, and Daily Fluency
Best AI Roleplay Chat Apps for Language Practice, Interview Prep, and Daily Fluency

2026 Latest Best Top-rated AI Roleplay Chat Apps for Language Practice, Interview Prep, and Daily Fluency! XIX.AI curates a powerful game-changing collection that offers free vs paid comparison, real-world tests, and updated rankings weekly. These must-try tools help you boost writing skills, overcome fluency challenges, and improve communication efficiency across all daily scenarios. Explore now to discover your perfect tool for language growth!

10 tools
xix.ai
Music composition AI Stem Separation Tools for Remix Production, Sampling Prep, and Karaoke Masters
AI Stem Separation Tools for Remix Production, Sampling Prep, and Karaoke Masters

2026 Latest Best Top-rated AI Stem Separation Tools Curated for Remix Production, Sampling Prep, and Karaoke Masters. These powerful game-changing tools offer real-world tests to deliver precise audio isolation, boosting productivity significantly. XIX.AI provides a weekly updated free vs paid comparison guide to help you find the must-try solution that suits your needs. Explore now to unlock your AI edge.

8 tools
xix.ai
Data Analysis AI SQL Copilots for Revenue Dashboards, Funnel Analysis, and Product Metrics
AI SQL Copilots for Revenue Dashboards, Funnel Analysis, and Product Metrics

2026 Latest Best AI SQL Copilots Ranked Top-Rated! XIX.AI curates a powerful game-changing collection for weekly updated real-world tests. These must-try tools help you generate accurate revenue dashboards, analyze sales funnels, and track product metrics swiftly, boosting productivity massively. Explore now to Discover your perfect tool for data-driven decision making! 238 characters

9 tools
xix.ai
Music composition Best AI Melody Writing Tools for Song Drafts
Best AI Melody Writing Tools for Song Drafts

2026 Latest Best Top-Rated AI Melody Writing Tools for Song Drafts! XIX.AI has curated a highly powerful game-changing collection that goes through rigorous real-world tests to deliver the best writing experience. You can find detailed free vs paid comparisons, accurate rankings, and must-try options designed to help you create stunning song drafts effortlessly and boost your creative productivity significantly. Explore now to discover your perfect tool!

8 tools
xix.ai
Comments (0)
0/500
OR