Jeeves. Reasoning improves Jev-like decision models

(github.com)

111 points | by nicowaltz 3 hours ago

17 comments

  • sharih 1 hour ago
    What is the point of this, if it is p90 17 seconds? Might as well use an LLM. The beauty of Jev is that it is dirt cheap and insanely fast.
    • zihotki 59 minutes ago
      I would hold your horses to paint it as dirt cheap.. In my cases for spam detection Luna was 20% cheaper due to prompt caching, although not as fast.
      • atombender 19 minutes ago
        > hold your horses to paint it as dirt cheap

        For a moment I thought this was going to be a metaphor — maybe an ancient Chinese proverb about how paint brushes are made from horsehair and how you can't hold the horse to paint before you've turned the hair into a brush.

      • tyre 11 minutes ago
        What are the costs compared to an ML model?
      • olgava 30 minutes ago
        [dead]
    • esafak 56 minutes ago
      Jev ought to offer a flex mode that uses their spare capacity for a discount.
    • mabini 53 minutes ago
      [dead]
  • thm 1 hour ago
    Ask Jeeves - Only took us 30 years to come full circle.
    • victordmor 6 minutes ago
      I met one of the founders once in Oakland. Amazing fella.
    • tmnstr85 1 hour ago
      this was the comment i came here for
      • onaclov2000 1 hour ago
        My bots are all named Jeeves lol. I have a CLI tool I use that connects up to a LLM I made and I call it Jeeves too ...so funny. I really didn't use Jeeves all that much I tended to use...I think it was called Web crawler pre-google era
  • TN1ck 35 minutes ago
    I just did a run with a benchmark I just used to test other models against. (It's about detecting irony in german soccer tweets). On my M5 Pro with 48GB it took over 30min to decide on just 100 tweets, the thinking definitely takes long.

    It performed quite below Jev, but above other open decision models I tested (68 correct vs 79 correct for Jev - see [1]). I'm running it for the moderation benchmark as well, but that will probably take a few hours on my machine.

    [1] https://tn1ck.com/blog/jevdit

  • loclol101 18 minutes ago
    How general really are these jev type models? Has anyone done any broad very cross-domain eval on them?
  • RamblingCTO 1 hour ago
    Super dope. If it would ship as prod ready code supporting mps as well that would be even doper.

    But funny that jev is getting its lunch eaten apparently in under two weeks?

    • danieltanfh95 19 minutes ago
      it just a classifier. I guess we have to thank typesafe for spending VC money on marketing classifiers as decision models instead.
    • pavlov 1 hour ago
      It’s ok, one week of AI hype is now enough to close a billion-dollar term sheet with VCs.
  • swader999 1 hour ago
    Seems like this is the way, a hybrid approach where some of the pipeline will be jev like and some traditional LLM depending on the nature of the work.
  • zerop 2 hours ago
    Are there "good" Open source Decision models built on Gemma-4 and also trainiable on own data?
  • alienbaby 2 hours ago
    Just curious, where has this term 'noul' come from for yes/no ansers?

    /a bit more digging and..

    A Noul performs a Bernoulli trial—an experiment with exactly two outcomes (yes or no)—but instead of picking one, it returns the calibrated probability (ranging from 0.0 to 1.0) that the statement is true.

    I hate it :)

    • LudwigNagasena 55 minutes ago
      In Bayesian statistics that’s called credence. Weird that they felt the need to invent a new term.
    • doginasuit 1 hour ago
      I like it. It is short and distinct which is a good fit for a primitive. It describes its fundamental meaning and draws a connotation with Boolean.
    • k__ 1 hour ago
      The whole "no hallucinations" premise is based on that.

      Like, yeah, you don't hallucinate, but only because you force the user to decide in the end.

      • kjs3 1 hour ago
        force the user to decide in the end

        And that's...bad?

      • doginasuit 1 hour ago
        That seems like the only possible way to eliminate hallucination, short of a model that is never wrong.
      • rusk 1 hour ago
        Wait til you hear about how digital circuits work at die level
    • keepitwiel 2 hours ago
      Bernoulli
    • user3939382 2 hours ago
      If you want to get super pedantic about what’s happening in a transistor every digital Boolean is actually this
      • kevindamm 2 hours ago
        Not quite.. that boolean is about whether the voltage exceeds some threshold. It's not about how close the voltage is to the circuit's maximum possible threshold, or how much it exceeds the threshold.

        In an analog circuit, maybe.

  • Naitik88 1 hour ago
    what about benchmark against smaller or bigger models? 9B looks too small for llm-level decisions.
  • woadwarrior01 1 hour ago
    This isn't really surprising. LLM reasoning and before that, chain of thought prompting are essentially forms of test-time compute scaling.
  • mxkuzn 1 hour ago
    interesting bench list, what about benchmark against smaller or bigger models? 9B looks too huge for small like laya, and too small for llm-level decisions.
  • captainbland 1 hour ago
    See if it can beat Jev's Pokémon benchmark
  • esafak 51 minutes ago
    Jev-like models give calibrated decision probabilities, but at low accuracy.

    So why didn't they show both??

  • AnodicElegy 1 hour ago
    I'm surprised we haven't seen a "Jehovah" yet.
    • jadar 1 hour ago
      With the amount of talk about "inventing god", I'm surprised too.
  • phplovesong 1 hour ago
    So "askjeeves" has been resurrected?
  • raverbashing 2 hours ago
    Jeeves, that's a name I haven't heard in a long time...
    • gizajob 1 hour ago
      Personally I’m happy that after a 30 year effort and hundreds of billions spent, AskJeeves finally works as intended.
    • fishfasell 2 hours ago
      If Jeeves returned as an AI chat bot it would be the most brilliant resurgence of nostalgia
      • grokkedit 1 hour ago
        jeeves is currently the name of my local hosted assistant, in its context there are rules that tell it to behave like good old jeeves.

        soon I'll make sure that my home assistant pod answers to "Hey jeeves"

    • kjs3 1 hour ago
      We locked him in the basement with Clippy, Bob and BonziBuddy. Who opened the damn basement door???
    • lherron 1 hour ago
      …a long time.
  • hjun1052 2 hours ago
    If the model does autoregressive reasoning before the decision, doesn't that give up much of what a Jev-style model buys you (a single forward pass, cheap calibrated probabilities)? Or is the point mainly to keep the typed output and probability interface while getting better accuracy on harder cases?