Alloy Compute Splits AI Models Across GPUs And FPGAs For Faster Replies

Alloy Compute claims sub-200-millisecond first-token responses from its combined AMD GPU and FPGA design, with performance data still unpublished. (Image: Shutterstock)
Alloy Compute claims sub-200-millisecond first-token responses from its combined AMD GPU and FPGA design, with performance data still unpublished. (Image: Shutterstock)

Alloy Compute says it has built an AI inference system that splits language model work between AMD graphics chips and programmable accelerators to cut response times.

Alloy Compute Splits Inference Workloads

The company described the design on Jul. 31 as a disaggregated architecture, one that sends each phase of a model's execution to the processor best suited for it. Most inference systems run every stage on a single chip type. AMD graphics processors take prompt prefill and attention, while FPGA accelerators handle token decoding and mixture-of-experts layers.

Alloy Compute has not published benchmark results and said it is still integrating the system while measuring latency, throughput, energy use and traffic between the two chips. Customer evaluation programs have not opened.

Also Read: Bitcoin ETFs Absorb $233.1M As A Single Fund Supplies Most Of It

Nour De Vos Backs Latency Guarantee

Founder and Chief Executive Nour de Vos said prefill, attention and token generation make different demands on hardware, so running a whole model on one architecture wastes capacity. The company said it is prepared to guarantee first-token responses under 200 milliseconds, though only for deployments that stay inside set limits on models, input length and concurrency. Early work covers quantized Qwen models.

De Vos came to AI infrastructure from large-scale cryptocurrency mining, where profits hinged on keeping machines busy and power bills low. He has said that same system-level thinking, treating chips, memory, networking and software as one unit, shaped this design.

Read Next: Polymarket Traders Give Spider-Man A 91% Shot At A Historic Debut

Murtuza Merchant profile photo

Murtuza Merchant

Murtuza is a seasoned finance journalist with extensive experience covering cryptocurrencies and blockchain technology. He has contributed to Benzinga and Cointelegraph, among other publications, reporting on emerging trends, the regulatory landscape, and more. Find him at @murtuza_merc on Twitter and mmerchant001 on Telegram. Disclosure: Murtuza holds ATOM, AKT, TIA, INJ, and OSMO.

Disclaimer and Risk Warning: The information provided in this article is for educational and informational purposes only and is based on the author's opinion. It does not constitute financial, investment, legal, or tax advice. Cryptocurrency assets are highly volatile and subject to high risk, including the risk of losing all or a substantial amount of your investment. Trading or holding crypto assets may not be suitable for all investors. The views expressed in this article are solely those of the author(s) and do not represent the official policy or position of Yellow, its founders, or its executives. Always conduct your own thorough research (D.Y.O.R.) and consult a licensed financial professional before making any investment decision.
Latest News
Show All News
Alloy Compute Splits AI Models Across GPUs And FPGAs For Faster Replies | Yellow