10 comments

  • herf 1 minute ago
    I have two NVIDIA GPUs (16GB+16GB) here, and it detects them each twice (says I have 4 GPUs). But then, it says most models are too big (anything >8GB?) and seems to run only on one GPU (5070ti).

    Unfortunately even with my 5070ti, llama.cpp seems to be about 20-30% faster at decode, running as:

    set CUDA_VISIBLE_DEVICES=0 build\bin\Release\llama-server -hf google/gemma-4-12B-it-qat-q4_0-gguf -ngl 99 --no-mmproj-offload -mg 0 -c 262144 -fa on --host 0.0.0.0

  • kmike84 21 minutes ago
    This seems to be a good idea. However, beating llama.cpp on speed is a low bar :)

    I found it to be a good baseline, but at least on Mac there was always something way faster, and/or with better memory requirements - like you said, ds4, omlx, mtplx, etc. It seems if you use local LLMs for real, there is very little reason not to use one of the more optimized engines.

    3 main failure modes I observed in the engines:

    * Not using best available spec decoding

    * Using too much VRAM for KV cache (e.g. KV cache used to take almost nothing in ds4, but huge amount of VRAM on unsloth/llama.cpp for deepseek models)

    * Degraded performance at large context sizes - benchmarks at 4K or 32K are awesome, but at realistic 100-200K it's slower than some stupid baseline

    • anerli 3 minutes ago
      Yeah these are all things that we directly tackle!

      Spec decoding: Models in our catalog come assigned with an assigned drafter model for speculative decoding based on the best known method and model available for that target model (support DFlash, DSpark, and DFlash2).

      Using too much memory for KV cache: We use a TurboQuant-inspired quantization of KV cache to 8-bit keys and 4-bit values. This drops KV memory usage by over half and also speeds up decode. Based on long context quality benchmarking we've done it does not seem to negatively impact retrieval or coherence over long context.

      Large context sizes: our KV quantization helps a lot for this, and we focus our optimizations on specifically longer-context requests since that's what most agent inference actually looks like.

  • lxe 15 minutes ago
    On my local inference box I have a perpetual codex thread open in my llama.cpp checkout that I periodically ask to take a look at currently pending llama.cpp PRs, do some research on latest MTP, Dflash and other prediction or attention optimizations, do research on the latest model quants and finetunes, take a look at localLlama Reddit threads and just do essentially a sweep of the frontier.

    Then it rebuilds latest llama.cpp, grabs the PRs it finds relevant to test against, and then it performs a benchmark and finalizes the upgrade and verifies what model, variant, or even a separate finetune that we should be running.

    Occasionally, it performs its own optimizations and commits, which then gets superseded by pull requests and merged code that essentially validates the model's own optimization directionality.

  • msdz 13 minutes ago
    Congratulations on the launch, it looks like an impressive product and tool!

    Q: From my (very, very limited!) understanding, I’m under the impression that part of the “inference engine inertia” is model- or at least architecture-specific code for most, if not each new open-weight model coming out. Assuming I got that right, do you plan on supporting everything vLLM/llama.cpp can do, such that Magnitude becomes a drop-in replacement for as many (economically/pareto-viable) models as possible, or do you want to focus on the best possible support for only a select few models/classes of models?

  • nateb2022 36 minutes ago
    Any source on the benchmarks/methodology besides the image? There's a ton of variance possible in llama.cpp's performance depending on how it was configured. I'd also like to see benchmarks against MLX.
    • anerli 10 minutes ago
      The benchmark we cited here is a simple prose-repetition task. We put the content of Moby Dick up to 64k context in the request, and then ask it to repeat the last section.

      For llama.cpp, we try to make the comparison as fair as possible by using similar settings. No speculative decoding, default prefill batch sizes, flash attention on.

      We tried also quantizing the KV cache to 8-bit keys and 4-bit values like we do in Magnitude, but this bombed decode speed for llama.cpp in our testing. Since it seems llama.cpp did not optimize that path, we used 16-bit KV instead.

      The source for the benchmark is available here also: https://github.com/magnitudedev/magnitude/tree/main/inferenc...

    • anerli 8 minutes ago
      Compared to MLX - we've done some rough benchmarking and we are outperforming any of the MLX-based engines we've compared to so far. Going to do more in depth benchmarking and release it soon.
  • sgtwompwomp 36 minutes ago
    This is dope, is this kind of like Wafer.ai but for local models? As in a coding agent optimizes the kernels so the local model runs continuously better? Cause that is compelling if so. If it’s more simple that’s cool too
  • kenzic 33 minutes ago
    How long does tuning take (on an M3 MacBook Pro for example)?
    • anerli 18 minutes ago
      Tuning is a one-time process that takes around ~1 minute whenever you download a new model. This is generally enough time to tune all the kernels' parameters to the point where tuning any longer asymptotes. Time can vary a little based on the hardware though.
      • kenzic 9 minutes ago
        Wow, that's impressive.
  • amirhesham 37 minutes ago
    Oh this is so cool. Curious about the business model, too.
  • p-e-w 44 minutes ago
    What is the business model?
    • anerli 21 minutes ago
      We envision a future where workloads are hybrid. Average consumer hardware will be able to handle a lot with local models, but you’ll still want to use cloud models for harder tasks. Magnitude will make it seamless to switch between the two, even for the same tasks (without breaking your prefix cache). We’ll charge per token for our inference cloud, using the same efficiencies we unlock for local inference to pass the savings on to you.
  • yolandac 29 minutes ago
    does it allow us to run larger models that weren't possible before?