TECHNOLOGY

Y Combinator Startup Launches Magnitude, an Open Source AI Engine That Tunes Itself to Your Hardware

Magnitude, a new open source inference engine from a YC S25 startup, compiles and tunes its kernels on the user's own device, claiming up to 2x faster performance than llama.cpp on local hardware.

Laptop on a desk running a local AI agent interface, with GPU hardware visible nearby in a home office settingTECHNOLOGY

Image: Magnitude · Uploaded by IntraGoals — usage rights confirmed

A Y Combinator-backed startup has launched Magnitude, an open source inference engine designed to run AI agents locally and automatically optimize itself for whatever hardware it is installed on. The project, unveiled in a "Launch HN" post on Hacker News, positions itself as a faster alternative to existing local inference tools such as llama.cpp, Ollama, and LM Studio.

The core pitch is speed through personalization. Rather than shipping generic kernels precompiled for broad categories of hardware, Magnitude compiles and tunes its kernels directly on a user's device before a model runs, tailoring performance to the exact chip in the machine. The company says this approach yields decode speeds up to 92 percent faster than llama.cpp on Apple's Metal framework and 19 percent faster on Nvidia's CUDA platform, with prefill speed gains of 9 percent and 23 percent respectively.

Magnitude ships as a desktop application for macOS, Windows, and Linux, with the command-line interface bundled in so no separate installation is required. The company says it runs on Apple Silicon, Nvidia and AMD GPUs, or CPU-only machines, with no fixed hardware minimum — smaller machines simply run smaller models, while more memory allows for larger ones.

Beyond raw speed, the company highlights several other efficiency features, including a 27 percent reduction in memory use per agent, memory that is freed automatically once an agent stops running, and shared prefix caches that keep concurrent sessions from slowing each other down. The engine also includes hand-optimized kernels for popular open-weight model families, which the company says is central to its performance edge over general-purpose engines.

Magnitude is built to plug directly into agents developers already use, with one-click connections for tools including Pi, OpenCode, Hermes, Codex, OpenClaw, Claude Code, Oh My Pi, and Cline. Any other tool can connect through an OpenAI-compatible API. The company emphasizes that the software is free, private, and released under the Apache 2.0 license: prompts, files, and models remain on the user's machine, and no internet connection is required once a model has been downloaded.

As is typical for a self-published product launch shared on Hacker News, the performance claims and technical specifications described here come from the company itself. Independent verification of the benchmark figures and broader claims was not available at the time of writing, and readers evaluating the tool for production use may want to consult the project's published benchmarks and documentation directly.

Sources and further readingGitHub - magnitudedev/magnitude: Open source inference engine for agents that optimizes itself for your exact hardware. Compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. Works on Apple Silicon, NVIDIA, AMD, or nothing but a CPU. ↗
ABOUT THE DESK

IntraGoals News Desk

IntraGoals reports on important changes in technology and work. We check each story for clear writing, trusted sources and useful information before it is published.

KEEP READING

Latest from IntraGoals.

All latest stories ↗