Hacker Newsnew | past | comments | ask | show | jobs | submit | nubg's commentslogin

Thank you, this benchmark to me proves that closed weight model companies are dangerous for our democracy and put kids at risk. They must be outlawed and all models must be made open weights!


welp, damning indictment. not sure if that means DS is super crap, or qwen is super good


Neither. Performance of all models is incredibly spikey.


if by "bot" we mean "not economically interesting", aren't websites within their right to filter our hacker-nerds who probably won't click on ads and buy useless objects?


You're absolutely right. The modern meaning of "bot" is "not economically exploitable". I wish they'd just say that rather than using the plausible deniability for what they doing re: "bots".


the ratio of this comment vs. parent is absurd. tfa's prior mentions "anti-cheat" in games. while cheaters do suck, a lot of draconian trash has crept in under the guise of that banner.

i have been finding my locked down experience has been becoming degraded lately fwiw. i am not economically "interesting", and the consequences of valuing privacy in 2026 is limiting information access. shame really.


As much as I want local and open-weights models to succeed, nothing beats a paid frontier model for now. Anybody who claims otherwise is simply not a daily user of such models. So this "investor" here should invest sime time in actually using the various LLM models and get a real taste of what it's like.


Their point isn’t that local models are better or even as good more but that if you can do 50%+ of tasks with local then that’s 50% of tokens that aren’t captured as compute done in data centers.


> As you can see, on average, SLMs are as good if not better than LLMs in 81.2% of the cases, with the LLMs having a significant advantage only in areas like engineering, life sciences, transportation and computer sciences.

So what I’m reading here is “LLMs have a significant advantage” in the most critical areas that have practically infinite demand for more intelligence.


The article is kinda dumb, and yes this is clearly the area where frontier models having and advantage matters the most, but I'd point out that these smaller open-weight models are performing better than the big Frontier models of just 4-6 months ago.

This means that the Frontier labs are under immense pressure to maintain that lead, and could end up in serious trouble if they stumble at all.

The other thing id point out is that a lot of us who are token-sensitive do things like build plans using expensive, smart models, and then execute those plans using cheaper dumber models.

Then there's the fact that we are still in the age of heavily subsidized Frontier subscriptions + tokenmaxxing initiatives from megacorps. Neither of which are sustainable, and will drive more usage to smaller open models once they end.


While that's true, the open / local models are getting good enough. Given time and the technology trend people may prefer a private local model for most use cases. Nobody is arguing that a Ferrari isn't a faster car, but the Honda is the more practical choice.


You completely missed the thesis here, and that is supported by the numbers being presented. It is that a large share of ordinary inference can be routed away from the hyperscalers.


How does a current local model compare to the best frontier model 12 months ago. Or 24 months ago?


It beats a frontier model from 12 months according to this bench: https://news.ycombinator.com/item?id=49334544

It is not the whole story, and knowledge is very lacking, but it has gotten a lot of attention. That model together with DeepSeek V4 Flash are the highlights of this summer on the open/local models side.


I have found qwen 3.8's coding quality using opencode to be similar to claude or gpt from 6-9 months ago, except much slower.


The authors of this article are ChatGPT and Claude.


> Ctrl + F

> "not"

> 15 results

ai;dr


why plus margin? the floor is cost


Floor is below cost if you think you can put one of your competitors out of business before you run out of money, and you'll make the money back after doing so. And this includes even after fines.


that's true, also see "loss leaders" like milk in supermarkets


very happy for you that things turned out this way


either the judge made the decision (in which case he has immunity) or he didn't (in which case he isn't the right person to sue)

the correct steps are appeals in the merit and disciplinary action against the judge


quantization level?


Its egregious the quant level isnt disclosed along the "Qwen" string. Everyone knows theres huge difference in speed/quality along the quant axis, I now attribute the ommision of such to deliberate choice to not curb the hype of the tittle.


they talk about the quants they tried in the article and settle on a Qwen3.8-27B-Hermes-iMatrix-NVFP4-Balanced.gguf which they calibrated on their own session traces and they pulled in 5 different llama.cpp pull requests to their local llama-server.


5.01 BPW custom hybrid: bulk NVFP4, selected Q5_K/Q6_K tensors from an iMatrix, Q6_K embeddings and Q8_0 lm_head. The iMatrix was built from 5,472 messages across 296 real Hermes sessions


Glad to read theyre not 296 fake Hermes sessions /s


The article mentions Q4, Q5, Q8, and NVFP4. It's total AI slop though, tough read.

In my testing I got 150 tokens/sec with a single 5090 RTX.


yeah, I was confused during the whole thing, i get 70 t/s on a 3090, which evens out around 50 t/s at 128k+ , have been running 3.6 and now 3.8 (both iq4_nl at 256k q4 kv) on the 3090 for months. I am confused as to what we 'discovered' here, it's a common config. and at less than 1/2 the price of the gpu (and double the bandwidth, though no fp4 cores to be fair).


A 3090 has 936 GB/s of bandwidth vs 432 GB/s on the 70W RTX PRO 4000 SFF, so 70 t/s there is not surprising. The interesting constraint here was fitting a workload-tuned 5.01 BPW quant + 256K + MTP into 24 GB while working with less than half the memory bandwidth


Which model/quant/command line did you use? I can barely get 100 token/secs and for sure clearly not a full context. With vllm, I am limited to 130k tokens with vllm + nvfp4.


If you've got RTX 5090, maybe try ninfer (https://github.com/Neroued/ninfer). Folks over on /r/localllama have been reporting wild prefill/token gen speeds with ninfer (NVFP4; 256k ctx).


Looks like slope benchmarks and results, and as usually, people are mixing MTP numbers with non MTP numbers. Or just 100 token input benchmarks. Or just failed ones as actual measures.

https://github.com/Neroued/ninfer/blob/master/docs/performan...

Category MTP3 stochastic sampler DFlash stochastic sampler DFlash greedy Code 1/15 natural stops; 0/15 prompt-complete 2/15 natural stops; 0/15 prompt-complete 0/15 natural stops Story 9/15 natural stops; the nine Chinese outputs pass requested division and minimum length 8/15 natural stops; the eight Chinese outputs pass requested division and minimum length 10/15 natural stops; five Chinese dialogue outputs are under length Translation 15/15 natural stops; 15/15 pass structural checks 15/15 natural stops; 15/15 pass structural checks 15/15 natural stops; 15/15 pass structural checks Structured 0/15 satisfy the requested complete record/script contract 0/15 satisfy the requested complete record/script contract 0/15 satisfy the requested complete record/script contract

And on my "own" "quick" benchmark, it's slower than vllm.


I don't have a 5090, so I can't really comment, but here's the relevant reddit thread from today where they report the numbers (including ninfer ones), and where you can make your case: https://www.reddit.com/r/LocalLLaMA/comments/1vqjeub/how_man...


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: