Strong disagree with the Anthropic being good at all part. This is not defending anyone else, but…
Anthropic leadership repeatedly presents themselves as uniquely morally qualified to steward agi and decide how humanity should get access to it. Yet they have repeatedly failed basic morality tests.
Pirating books for financial gain. The newer Sony/Warner music case shows this is pattern behavior.
Aggressively scraping other people's works, despite the authors' requests not to do so.
Then applying massive usage restrictions on their own work.
And probably the most disqualifying is backing away from their own hard AI safety commitments.
Which safety commitments did they back away from? My understanding is that they believe safety can only be researched from the frontier, and so they're trying to be pragmatic to stay near the frontier (and viable) in their choices.
From what I know, the "books3" dataset was normalised in the LLM and research ecosystem, where collected datasets were seen as valid to train on and/or fair use. I'm not sure any of the major frontier companies are free from that, if we don't believe it was fair use.
I do think most of their choices are explainable by "they just believe in agi risk". You truly wouldn't want non-agi-pilled companies to train on your data and approach the frontier if you were worried. You might slightly hurt your own business with safety filters (that no one else does) if you were worried. They are less worried about other "moral" decisions like "sharing" if they conflict with AGI: the research they still share is all of their safety research.
This definitely doesn't make them "good", but they do seem fairly "consistent". Most of these issues were talked about publicly by the founders long before Anthropic was founded and/or the AI race+money appeared.
as a safety commitment they walked away from - they were similarly negligent to openai in terms of asking a model with a hacking based harness to go have fun, and then not watching it at all while it could do harmful and illegal stuff.
thats not something you expect from a company that "believes in agi risk"
I don't think "not watching it at all" is completely fair. They thought they had sandboxing/monitoring etc. I definitely won't say they're free of mistakes though.
Note that the companies that haven't faced these issues so far are the ones that don't do safety testing, or don't have frontier models. I'm not sure who I would pick as "better" on any of this right now.
I will give Anthropic credit for standing up against the department of war. The bar is incredibly low, but not doing domestic surveillance and not creating autonomous weapons are laudable.
That doesn’t mean I like them pirating books and being shady about tokens and paternalistic “safety”
It's hard for me to see much difference between Amodei and Sama. My guess is they're both savvy SV CEOs who will bend their message, alliances and principles pretty far if that's what it takes to get ahead. Musk and Zuck feel like something else entirely, with all the reactionary imagery, populist bullshit and the societal damage around their platforms.
It's almost like running a trillion-dollar business with neck-to-neck competition against other frontier labs and even state-sponsored efforts requires some ethical trade-off.
Pirating books is just straight up morally correct. I don't like Anthropic's bullshit "safety" filters, but training on shadow library data? Yeah no, it makes sense.
It makes a lot more sense than having to work around copyright by scanning out physical books. Unfortunately, one was ruled legal and the other was not.
> This means that when this ram hits EOL it -won't- flood the consumer market in a meaningful way. It is disposed of because most people won't put a server GPU into anything. It is like a used big rig, built for purpose and when it reaches end of life it is truly at end of life.
This was true until the last year.
P40, P100, V100, A40, etc are selling like hotcakes and prices have been rapidly increasing.
Not sure about that, V100s are all selling for less than $1k despite being 24-32GB VRAM around 1TB/s memory bandwidth, which is the same price point that 4090s and 5090s are commanding $4k-$5k for. Hardly hotcakes.
Please don’t buy a DGX Spark unless all three of these are true:
- You value simplicity more than performance or price-to-performance.
- You accept that the hardware will depreciate rapidly.
- You’re prepared to buy two or four of them.
OR:
- You want to run frontier models right now as cheaply as possible
- You want to run high-parameter models on a 15a breaker/line
Otherwise, get a normal, high-bandwidth GPU.
A single Spark gives you roughly 115 GB of usable memory compared with the 24–32 GB found on many lower-cost GPUs. It's certainly a big increase, but in practice it does not unlock dramatically better models.
- One Spark: More memory, but mostly enough for poor-quality, extremely low-bit quants of larger models.
- Two Sparks: Enough for mid-tier parameter models at reasonable quants, such as DeepSeek V4 Flash and HY3.
- Four Sparks: Enough for GLM 5.2 at a reasonable quant. You'll need a $1000+ switch too.
The problem is that Sparks are slow compared with almost everything else in their price range. Many factors affect inference speed, but memory bandwidth is one of the biggest. A $4,000-plus DGX Spark provides only 273 GB/s.
Yes, the Spark has substantially more memory. But going from roughly 24 GB to 115 GB does not necessarily unlock substantially better model quality. In many cases, it only lets you load heavily compressed 2-bit versions of larger models, such as DeepSeek V4 Flash, with serious quality degradation.
24–32 GB is currently a sweet spot. Models such as Qwen 3.6 27B and 35B-A3B:
- Perform far above what their parameter counts suggest.
- Fit comfortably within 24–32 GB of VRAM at reasonable quantization levels.
A 4-bit quant of Qwen 3.6 27b (18 GB) will out-perform a 2-bit quant of DeepSeek v4 Flash (90gb).
Instead of the Spark, if I had a roughly $4,000 budget...
Assuming I already had a reasonably modern desktop:
- One RTX 5090, RTX 5000 Pro, or RTX 4500 Pro.
- Two RTX 3090s, RTX 4000 Pros, or R9700s, provided the motherboard can bifurcate two physical x16 slots into x8/x8.
If I were building a system from scratch:
- A DDR4- or PCIe 4.0-era consumer CPU and motherboard that supports x8/x8 bifurcation.
- Two RTX 3090s, RTX 4000 Pros, or R9700s.
If I were already planning to buy a new Mac:
- A MacBook Pro M5 with 64 GB or 128 GB of unified memory.
For context, these are the systems I currently run:
- EPYC Turin with four RTX 6000 Pro Max-Qs.
- EPYC Milan with four RTX 3090s.
- AM4 with two RTX 3090s.
- AM4 with two RTX 3090s.
- Intel Raptor Lake with two RTX 5060 Ti.
- MacBook Pro M3 128GB Unified
I agree with this. I have a dual rtx4090 machine, a 128gb m5 max mbp, and a dgx spark. RedHatAI/gemma-4-26B-A4B-it-FP8-dynamic on the dual rtx4090 machine under vLLM absolutely slays at token speed. I'm frustrated with how hard it is to realize goodness on the dgx spark.
I'm actually confused by the post - indeed why choosing the 27B dense and the 35B MoE models? Given the structure of the post I honestly have the impression that the Spark was purchased out of curiosity more than for a specific purpose (LLMs).
Ah yeah, your reaction makes sense. In the circles I run in, there is a lot of hype around Sparks for inference, so my gut reaction is to respond with this type of warning.
I did not intend to imply that the post author was advocating that they're great for inference, as they're obviously not.
Actually, its ~119GB usable (I have one). You can shutdown a lot of unneeded services if you only use it remote which will trim a lot more fat. You can enable the RDP service if you don't want to sit in front of it but get a desktop interface.
The "shitty" network is 10Gb/s and wifi7! You get a twin QSFP28DDlol+++ (I jest) that each run at 200Gb/s - not for the casual home user but handy at work, although I "only" have 40Gb/s on my switches sigh. With and no switch two you can do a three node cluster with some careful networking. If you want to do more then a switch is needed and it will need to be pretty funky! That said you could wire them up in a circle and use VLANs and MSTP and accept less than 200Gb/s per link. You'll probably need Openvswitch and a lie down afterwards.
I'm not a fan of the Gnome desktop but it works well enough and I think the Nvidia customised Ubuntu is well thought out. You get all the complicated NVidia extras pre-installed, along with docker (full fat, not the Ubuntu one) for a fairly quick start. It includes Ubuntu Pro which is free for five systems anyway but its nice to see it pre-installed.
We blew abut £4000 on one and it will pay for itself in a few months. I tried pricing up an Apple thingie and the Store wouldn't offer me more than 96Gb of RAM and a delivery date in Q3 at the earliest. Our Spark rocked up next day. They seem to come in 1TB or 4TB SSD variants. 1TB is enough for me and saves a lot of cash - keep an eye on your model downloads and ruthlessly delete old experiments. docker system prune.
We went for the Asus variant that has active cooling and I stuck it in the ceiling cable tray over our computer room racks. It sits on 1½" stainless steel mesh with lots of clearance in an actively cooled environment.
That ~119 GB optimization is great. Hadn't seen that yet. That pops a 4x-Spark setup to 476 GB, which gives a lot more headroom for running something like GLM 5.2 at 4-bit, which isn't completely terrible.
I should have prefaced my post - I almost bought four Sparks a couple months ago, but ultimately opted to buy two more RTX 6000 Pro Max-Q's.
It was a painful choice because the two 6000's were more expensive than four Sparks, and ultimately gave me only 384 GB VRAM.
It was even more painful when GLM 5.2 was released, and a 4x Spark setup could run it at a decent quant, but 4x 6000's cannot with any headroom.
But the 6k's absolutely destroy the Sparks on prefill and inference speed. Model intelligence is compressing. The smaller VRAM pool will matter less over time than slower prefill/inference speed.
That is to say, I'm sure we'll end up with <500B parameter models that are Fable-level in the next 8 months or so. Performant quants of those will fit comfortably in 384 GB.
In your previous posting you said that for more than 2 Sparks you also need to buy a very expensive switch.
That is not really true. Once you have 2 fast Ethernet ports, like DGX Spark has, you can interconnect any number of systems without using a switch.
In the simplest case, you just daisy chain the systems and you configure in Linux the Ethernet interfaces as bridges.
For better performance, you can close the chain into a ring, with an extra cable. In this case the Linux configuration is more complex, because you must do IP-level routing, preferably with a routing protocol like OSPF, to be able to double the throughput between 2 systems, by using both paths through the ring.
The only advantage of a switch is a greater throughput when there are 4 or more systems, which happens only when all the interconnected systems are attempting to communicate simultaneously, in which case having only 2 paths through a ring will serialize some of the transmissions.
When only 2 systems attempt to communicate with each other, a ring has double throughput in comparison with a switch. You need 2 switches to match the throughput that a ring has, as long as it does not become congested. Or you could use one double-sized switch, with the ports partitioned between 2 VLANs, with each DGX Spark connected to 2 ports, in different VLANs. Both 2 switches or a double-sized switch would greatly increase the price.
For only 3 systems, a switch is useless, as it is worse than connecting the systems in a triangle and configuring the IP addresses for point-to-point links (the configuration as bridges is needed only for 4 or more daisy-chained systems). The configuration with 3 DGX Sparks seems optimal from the PoV of the performance per dollar ratio.
On each Spark you can bridge the two interfaces but you cannot bridge those bridges at layer two and it sounds from your post that you are confusing layer two and three (ARPA).
So if you have three Sparks, you have three two port switches. Take three two port switches and connect them in a ring and you have a collision storm. You have to use STP or similar to sort that out.
The ring is now collapsed to A-B-C (with no A-C) and to get from A to C you have to go via B which halves thoughput for those two paths: A-C and C-A.
Another "option" is to define two VLANs and use MSTP and layer three and hope the protocol that shuffles data can route packets effectively. So you get two sets of paths at layer two: A1-B1-C1-A1 and A2-B2-C2-A2 and you cut the two graphs at different points, so you get two disjoint graphs, for example: A1-B1-C1 and B2-C2-A2. However they will still overlap and it all falls apart, somewhat.
You can only link two Sparks together without a switch to get the full possible throughput. That is why Nvidia only offer a two pack option. Three or more requires a switch to avoid a path being used twice. Ethernet is not Token Ring! Mind you Tring was 16Mb/s - luxury when ethernet was 10Mb/s 8)
I've also had this discussion with folk when it comes to "hyper converged" virty systems: So we have three nodes with two NICs each for Ceph and we cable them up in a ring and it goes weird ... . At least there I'm only generally specifying 10G.
Now, do you really need the full 200Mb/s for clustering SParks? No idea yet, I've only got one and I've only got 40Gb/s QSFP+ to play with on my switches.
However I'm seeing value in the one box already at the moment and if that scales then I'll be buying more Sparks and a really funky switch when I need three Sparks and perhaps another switch down the line with some really fancy VLT or whatever wankery is the stacking du jour thing!
This is assigning intent without evidence, as is common in tribal politics. A non-charged assessment might use the phrase "abrupt cancelling."
We cannot create a better republic without constructive discourse, and we cannot have constructive discourse when we default to characterizing the views, concerns, and actions of those we disagree with as rooted in moral failure. Even if it is true from time to time.
This chiding from you would be better received if there was a shred of evidence that the other tribe is even slightly receptive to this kind of discourse.
Unfortunately, the norms of discourse are pretty much gone. This is terrifying in the long-term.
If you were to try to convince me a 2:1 immigrant to local birth ratio here in Australia is a net good for the country, you’re first going to have to convince me your a reasonable person to have a conversation with.
If you jump straight in with claims against me that I’m -ist and -ic that’s going to be more difficult.
Here you are again [0], unable to represent comments in good faith.
Nobody said "non-white" and it isn't even implied because a significant proportion -- 35-40% of Australian Permanent Residents (US green card equivalent) -- come from EU/"White" countries.
The suggestions above are consistent with my request for you to review the HN comment guidelines.
Complaining about 2:1 immigrant to local birth rate has absolutely nothing to do with long-time permanent residents (who are locals). It's clearly a fear that white culture is being overrun by brown/chinese people.
Your comments degrade the discourse at least as much as you think mine do.
> This is assigning intent without evidence, as is common in tribal politics
You are calling for constructive discourse and yet your response is an accusation of dishonesty. A non-charged assessment might use the phrase "without presenting evidence".
> Does anyone here have experience running large models in a multi-GPU setup with several RTX 6000s in a high-concurrency regime and with large context lengths? (something like Deepseek 4 Flash, Minimax 2.7 etc.)
What an entirely unserious company. So glad I dumped Claude Code last summer after being gaslit by Anthropic over service degrades. I was fine with the service degrades, totally understandable. Being lied to, not at all.
OpenAI and Altman present a whole set of different concerns, but Codex does not get in my way of doing what I want to at all. Also let me use pi without a banhammer.
Apple only makes disposable devices now. They're a megacorp can negotiate massive discounts at every stage of the supply chain.
I've helped several people in the last few years set up new Macs, replacing ones that were only 1-2 years old, because they ran out of storage.
Additionally, the comparison doesn't even hold true when you need more than the base configs from Apple, given their ridiculous upgrade pricing. I'm writing this on a $6,000USD M3 MBP with 128gb/4tb. It would have been substantially cheaper to build out on a Framework.
IMO it’s a reasonable point to make when compared to something like the Framework. And it took legal action to get them to offer battery replacements for iPhones, I don’t think you can really claim they’re passionate about component reuse.
> "[...] they just last way longer than any of their competitors."
Citations, please.
In the meantime, an anecdote: My oldest modern-era (64-bit) daily driver is one I use heavily since 15 years. An HP, 16 or 17 years old. The only component that ever caked-out in that bird was the original mechanical hard drive, which died just this year. Similar experiences with IBM/Lenovo, Panasonic and Fujitsu. Apple laptops I don't even look at for they don't offer anything I need.
>Everything points to commoditization of models. Open/distilled models lag behind frontier only by 6-12 months.
Yes, but every high performing open weights model coming out of China has (supposedly) been caught distilling frontier models.
It seems like a lot of people are making assumptions about the state of the open weights ecosystem based on information that may not be accurate. And if the big labs are able to reliably block distillation, we could see divergence between the two groups in terms of performance.
> And if the big labs are able to reliably block distillation,
The big labs will not be able to reliably block distillation without further inhibiting general use of the models, which itself will help tip the balance away from commercial models.
No, you're wrong. It won't tip it away from commercial models. Trying to run open weight modesl to do inference is something 99% of people around the world can't do because it's expensive and technically challenging and the results are poor compared to the main companies. If they get rid of free usage people will simply pay for it.
> Trying to run open weight modesl to do inference is something 99% of people around the world can't do because it's expensive and technically challenging and the results are poor compared to the main companies.
Just because a model is open doesn't mean that there aren't services that will run it for you (and which won't share any limits that the commercial model vendors impose to fight distillation because neither the host not the model creator cares if you are using the service to distill the model.)
Many users of, particularly the larger, open models now are using such services, not running them using their own local or cloud compute.
The article is obviously bad (I quitted reading after the second paragraph) but one side effect of AI training is the increasing cost of hardware. We have commoditization of models... while reversing commoditization of hardware.
Anthropic leadership repeatedly presents themselves as uniquely morally qualified to steward agi and decide how humanity should get access to it. Yet they have repeatedly failed basic morality tests.
Pirating books for financial gain. The newer Sony/Warner music case shows this is pattern behavior.
Aggressively scraping other people's works, despite the authors' requests not to do so.
Then applying massive usage restrictions on their own work.
And probably the most disqualifying is backing away from their own hard AI safety commitments.
reply