AI Infrastructure

Etched Raises $500M to Take On Nvidia With Transformer Chips

Etched raised about $500 million led by Stripes at a reported $5 billion valuation to take on Nvidia. Its TSMC 4nm Sohu ASIC runs transformer models from OpenAI, Google, Microsoft and Anthropic.

Etched Raises $500M to Take On Nvidia With Transformer Chips — article cover

On January 13, 2026, Bloomberg reported that Etched, a San Francisco AI chip startup, had raised roughly $500 million in a new round aimed squarely at NVIDIA’s dominance in AI accelerators. Data Center Dynamics, citing sources, reported the round was led by investment firm Stripes at a valuation of around $5 billion; Built In SF confirmed that Peter Thiel is among the investors.

The round extends an already steep capital trajectory: Etched closed a $120 million Series A in 2024. Going from $120 million raised to a $5 billion valuation in barely a year says plenty about how hungry capital markets are for any credible alternative to NVIDIA.

Sohu: An ASIC Built for Transformers and Nothing Else

Etched’s flagship product is a chip called Sohu. Its fundamental difference from a general-purpose GPU: the transformer architecture is baked directly into the silicon, making it an application-specific integrated circuit designed around one model architecture. Built In SF reports that Sohu is fabricated by TSMC on a 4nm process, is intended to be more efficient than general-purpose GPUs, and can run models from Google, Microsoft, OpenAI, and Anthropic.

The bet is easy to state. Nearly every large language model today is built on transformers, so the workload is remarkably homogeneous — spend every transistor on that one kind of computation and trade flexibility for efficiency. A general-purpose GPU has to cope with every imaginable workload: graphics lineage, sparse matrix code, scientific computing, whatever a customer brings. An ASIC designed for one architecture does not carry any of that overhead. At the same process node, that alone can open a gap on paper, and because the chip targets a narrower problem, more of the silicon budget goes to the arithmetic that transformer inference actually performs.

Why Now: The Inference Boom

Timing is the real backdrop of this round. Inference, not training, is becoming the dominant driver of AI compute demand: agentic applications generate dozens of model calls per task, and search, summarization, and coding agents all burn inference tokens continuously. NVIDIA’s capacity and pricing have become the industry’s chokepoint — waiting for allocations, rationing throughput, and watching bills climb has been the shared experience of platform teams since 2025.

The hyperscalers have already voted with their budgets: Google’s TPU and Amazon’s Trainium both follow the custom-ASIC path and are gradually opening up beyond their own clouds. Etched is effectively the same logic packaged as an independent company, aimed at AI platforms that cannot afford to design their own silicon. In the same week, Meta announced its Meta Compute program with a target of tens of gigawatts of infrastructure this decade — when upstream and downstream both escalate simultaneously, the market is treating inference-specific silicon as a core component of the next buildout cycle.

The Risk of Betting on One Architecture

The risks are just as concrete. First, architecture risk: the efficiency case rests entirely on transformers staying dominant. If frontier models shift toward hybrid architectures or a new paradigm, the flexibility penalty of an ASIC gets magnified overnight. Second, ecosystem risk: two decades of CUDA software stack and developer inertia do not convert into market share through hardware efficiency alone — every missing kernel library, framework integration, and deployment tool is homework a newcomer must finish. Third, execution risk: between an efficiency claim and mass deployment sit yields, supply chains, and the realities of production data centers.

For buyers, the sensible posture is to treat an ASIC like this as one leg of a multi-supplier compute strategy rather than a replacement: add a supply line optimized for transformer inference alongside general-purpose GPUs, and split workloads to hedge pricing and capacity risk on any single vendor. The evaluation playbook is equally concrete — pilot on a bounded inference workload, measure tokens per dollar and tokens per watt against the incumbent GPUs, and only then decide how much traffic to shift. That is the shared signal of early 2026’s capital flows — in the same week xAI closed its $20 billion Series E and Meta announced a gigawatt-scale infrastructure program, capital is doubling down on every layer of the compute supply chain.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

SHAREXEMAIL