French AI Startup ZML Launches Free Tool to Speed Up Models on Any Chip

ZML, a rising French AI company, has just released a free inference server called LLMD that boosts performance for large language models across many different chips.
The tool supports hardware from major players like Nvidia, AMD, Google, Apple and Intel, helping users avoid being locked into one vendor.
Why Multi-Chip AI Matters for Everyday Users
This release comes at a time when AI costs are rising fast, making efficient inference crucial for apps we use daily like chatbots and image generators.
By allowing a mix of chips, the software could cut energy use and expenses, which ultimately means faster and more affordable AI features for regular people.
Founder Steeve Morin, backed by investor Yann LeCun, aims to let companies build custom systems that run at peak speed without barriers.
Europe gains an edge here too, as the Paris-based team of 20 works with local chip innovators to co-design future hardware.
Broader Shifts in the Global AI Race
Unlike earlier open frameworks from the same company, this new product stays closed source for now but offers free access to gather real-world data.
With 20 million dollars raised, the startup is positioned to expand quickly into new releases that could reshape how data centers operate.
Looking ahead, wider adoption may push the industry toward hybrid setups that mix expensive and cheaper chips for better overall results.
For the average person, this development signals a future where AI feels more seamless and less tied to big tech monopolies.
Success will depend on how many users try the free version and what feedback shapes the paid options later.








