
Google is designing a new custom artificial-intelligence accelerator specifically optimized for its Gemini family of models, a move that would reduce the company's dependence on Nvidia and lower the cost of running some of the industry's largest neural networks.
The effort, which has been underway for several months inside Google's semiconductor division, targets the inference stage of AI workloads — the moment when a trained model answers a query, writes code, or generates an image. Inference now accounts for the majority of computing costs at leading AI labs, and Google's internal projections suggest that a custom chip could cut per-query expenses by 30 to 50 percent compared with off-the-shelf graphics processing units.
People familiar with the project say the chip borrows architectural ideas from Google's existing Tensor Processing Units but adds specialized memory bandwidth and matrix-math units tuned to Gemini's transformer architecture. The design emphasizes high throughput for long-context windows, one of Gemini's distinguishing features, while keeping power consumption within the strict thermal limits of existing data-center racks.
If successful, the chip would mark Google's most aggressive vertical-integration push since it began designing TPUs nearly a decade ago. It would also place the search giant in direct competition with a growing list of tech companies building their own AI silicon, including Meta, Microsoft, and Amazon, as well as with Nvidia, whose data-center revenue has become the single largest driver of the semiconductor industry's growth.
Analysts caution that custom chips require years to mature. Google's first TPU took three generations before it matched Nvidia GPUs on mainstream benchmarks, and the company has previously shelved consumer-chip projects such as the Tensor smartphone processor for cost reasons. Still, with Gemini now serving hundreds of millions of users across Search, Workspace, and Cloud, even modest efficiency gains could translate into billions of dollars in annual savings.
Google has not publicly confirmed the chip's existence or a timeline for deployment. Industry observers expect a limited internal rollout by late 2027, with broader availability for Google Cloud customers following in 2028 if initial performance targets are met.
Image source: i.ibb.co