The Hunt for AI Compute: Is Cerebras the Next Big Thing?

The surging demand for computing power to run AI models continues to intensify, yet industry players face two critical hurdles: securing the appropriate hardware and deploying it in data centers to begin generating revenue.
General Compute, a new inference-focused neocloud provider specializing in the deployment phase where models respond to users rather than undergoing training, offers solutions that highlight the evolving direction of the AI ecosystem. These strategic advantages facilitated a $15 million seed funding round at a $60 million post-money valuation, led by FUSE VC, with additional participation from Carya Venture Partners and Village Global Ventures.
Identifying the optimal chip is the first challenge. While GPU demand has skyrocketed, it is increasingly recognized that they are not the most efficient choice for running trained AI models. The computational requirements for inference differ significantly from those of training, prompting the development of a new class of chips designed specifically for this purpose. Nvidia’s $20 billion acquisition of Groq in December and Cerebras’ $57 billion IPO last week underscore this shift.
With capacity constraints at Groq and Cerebras, General Compute co-founders CEO Finn Puklowski and CTO Jason Goodison identified an alternative. They have partnered with SambaNova, an Intel-backed chipmaker specializing in inference solutions that has received less attention in Silicon Valley.
This dynamic may shift when SambaNova releases its new chips later this year. The architecture offers greater flexibility and utilizes more memory to store context during inference calculations. SambaNova claims these chips outperform not only GPUs but also specialized hardware from competitors like Groq and Cerebras. According to Puklowski, the new chips will achieve speeds of 600 to 700 tokens per second, compared to approximately 250 tokens per second for standard GPUs.
General Compute has placed an order for $300 million worth of SambaNova’s SN50 chips and aims to be the first neocloud to deploy them.
These chips also address the second major challenge: deployment location. Since they are air-cooled rather than water-cooled and consume less power, they can be integrated into existing data center facilities without requiring new infrastructure investments.
Puklowski is pursuing colocation agreements, where General Compute installs its hardware in third-party facilities. These partnerships extend beyond traditional data center providers to include crypto miners seeking to repurpose their infrastructure, particularly as the cost of mining Bitcoin often exceeds its market value.
General Compute launched its cloud service last week, claiming it currently offers the fastest performance for MiniMax 2.7, a powerful open-source large language model.
Joe Hasselmann, a venture investor who backed Groq in 2021 during the early stages of the inference boom, launched a new fund this year, Evercrest Capital Partners, focused on AI. He made General Compute his first investment. Hasselmann sees parallels between SambaNova’s partnership with General Compute and Coreweave’s relationship with Nvidia, as well as the combination of Groq’s chip manufacturing with its previous cloud offerings.
“They need a diverse customer base that will place their chips in high-growth environments,” Hasselmann noted. “Just as General Compute is betting on SambaNova, SambaNova is betting on General Compute.”
The key question remains which computer architecture will capture the most value in the future of AI. Inference clouds represent an implicit bet on a landscape with multiple models and agents, where no single provider dominates, and speed and cost become the primary competitive factors. This is reflected in OpenRouter’s recent $113 million Series B raise, which highlights its ability to provide customers access to multiple models to optimize token spending.
Speed is crucial for pricing and capability. Puklowski aims to reduce hour-long coding agent workloads to five or ten minutes and make audio agents for customer service more economical by enabling faster inference for effective conversation.
“If you use ChatGPT and it generates 50 tokens per second, that’s still much faster than human reading speed,” Puklowski told TechCrunch. “However, with the shift to agent-to-agent interactions, where agents read on our behalf or query databases, they need to operate at much higher speeds.”
Related article
The sameness problem behind those unappetizing AI-generated menus
At first, you might question your sanity. You step into a café and scan a menu filled with bagel sandwiches, yet every image appears unnervingly perfect—flawlessly symmetrical and overly smooth—triggering an instinctive sense that something is off. Y
Hello Robot Readies Home Robots for Silicon Valley
Martinez, California, sits on the far northeastern edge of the San Francisco Bay Area, a world away from Silicon Valley’s tech hub. Here lives Hello Robot, a startup deliberately distancing itself from the grandiose claims of its southern rivals, ope
Former Infosys Chief’s AI Startup Secures Another $53M
Hang Ten Systems, an AI startup established by former Infosys CEO Vishal Sikka just four months ago, has secured an additional $53 million in seed funding. This latest investment round was finalized merely five weeks after the initial $32 million see
Related Special Topic Recommendations
Comments (0)
0/500

The surging demand for computing power to run AI models continues to intensify, yet industry players face two critical hurdles: securing the appropriate hardware and deploying it in data centers to begin generating revenue.
General Compute, a new inference-focused neocloud provider specializing in the deployment phase where models respond to users rather than undergoing training, offers solutions that highlight the evolving direction of the AI ecosystem. These strategic advantages facilitated a $15 million seed funding round at a $60 million post-money valuation, led by FUSE VC, with additional participation from Carya Venture Partners and Village Global Ventures.
Identifying the optimal chip is the first challenge. While GPU demand has skyrocketed, it is increasingly recognized that they are not the most efficient choice for running trained AI models. The computational requirements for inference differ significantly from those of training, prompting the development of a new class of chips designed specifically for this purpose. Nvidia’s $20 billion acquisition of Groq in December and Cerebras’ $57 billion IPO last week underscore this shift.
With capacity constraints at Groq and Cerebras, General Compute co-founders CEO Finn Puklowski and CTO Jason Goodison identified an alternative. They have partnered with SambaNova, an Intel-backed chipmaker specializing in inference solutions that has received less attention in Silicon Valley.
This dynamic may shift when SambaNova releases its new chips later this year. The architecture offers greater flexibility and utilizes more memory to store context during inference calculations. SambaNova claims these chips outperform not only GPUs but also specialized hardware from competitors like Groq and Cerebras. According to Puklowski, the new chips will achieve speeds of 600 to 700 tokens per second, compared to approximately 250 tokens per second for standard GPUs.
General Compute has placed an order for $300 million worth of SambaNova’s SN50 chips and aims to be the first neocloud to deploy them.
These chips also address the second major challenge: deployment location. Since they are air-cooled rather than water-cooled and consume less power, they can be integrated into existing data center facilities without requiring new infrastructure investments.
Puklowski is pursuing colocation agreements, where General Compute installs its hardware in third-party facilities. These partnerships extend beyond traditional data center providers to include crypto miners seeking to repurpose their infrastructure, particularly as the cost of mining Bitcoin often exceeds its market value.
General Compute launched its cloud service last week, claiming it currently offers the fastest performance for MiniMax 2.7, a powerful open-source large language model.
Joe Hasselmann, a venture investor who backed Groq in 2021 during the early stages of the inference boom, launched a new fund this year, Evercrest Capital Partners, focused on AI. He made General Compute his first investment. Hasselmann sees parallels between SambaNova’s partnership with General Compute and Coreweave’s relationship with Nvidia, as well as the combination of Groq’s chip manufacturing with its previous cloud offerings.
“They need a diverse customer base that will place their chips in high-growth environments,” Hasselmann noted. “Just as General Compute is betting on SambaNova, SambaNova is betting on General Compute.”
The key question remains which computer architecture will capture the most value in the future of AI. Inference clouds represent an implicit bet on a landscape with multiple models and agents, where no single provider dominates, and speed and cost become the primary competitive factors. This is reflected in OpenRouter’s recent $113 million Series B raise, which highlights its ability to provide customers access to multiple models to optimize token spending.
Speed is crucial for pricing and capability. Puklowski aims to reduce hour-long coding agent workloads to five or ten minutes and make audio agents for customer service more economical by enabling faster inference for effective conversation.
“If you use ChatGPT and it generates 50 tokens per second, that’s still much faster than human reading speed,” Puklowski told TechCrunch. “However, with the shift to agent-to-agent interactions, where agents read on our behalf or query databases, they need to operate at much higher speeds.”
The sameness problem behind those unappetizing AI-generated menus
At first, you might question your sanity. You step into a café and scan a menu filled with bagel sandwiches, yet every image appears unnervingly perfect—flawlessly symmetrical and overly smooth—triggering an instinctive sense that something is off. Y
Hello Robot Readies Home Robots for Silicon Valley
Martinez, California, sits on the far northeastern edge of the San Francisco Bay Area, a world away from Silicon Valley’s tech hub. Here lives Hello Robot, a startup deliberately distancing itself from the grandiose claims of its southern rivals, ope
Former Infosys Chief’s AI Startup Secures Another $53M
Hang Ten Systems, an AI startup established by former Infosys CEO Vishal Sikka just four months ago, has secured an additional $53 million in seed funding. This latest investment round was finalized merely five weeks after the initial $32 million see





Home






