I believe Opus 5 isn't meant to be spoken to by humans. It's great at executing but I reckon it's intended to be spoken to by other models such as Fable. I use Fable as the orchestrator, only speak with Fable, and all implementation, recon, design etc happens with Opus 5, with Fable reviewing (and translating).
I've really gone in the opposite direction: having a dumber model orchestrate. In my case, it's usually a Luna orchestrator spawning Sol/Astra subagents to do the "big brain" work of planning and reviewing.
Reason I went with "dumb orchestrator" was just to save tokens. Having Opus/Sol (let alone Fable/Astra) orchestrate was burning tokens like crazy for me even when much of the gruntwork was being done by Luna/Sonnet/Haiku subagents. (Luna is also really good, like way better than Sonnet...) Perhaps it was a skill issue on my end though, maybe I wasn't just managing context properly.
It's amazing how quickly Fable went from 'Game-changing model that needs to be banned' to 'Yeah it's alright, but OpenAI is also just as good and there are a couple of good open weight alternatives that are equivalent for almost everything'
The hype cycles are shortening, perhaps we really are reaching some kind of plateau this time (famous last words)
Why are you using the product release date for Kimi K3 and the training date for Fable? Either use the release date for both (6 weeks apart) or if you have it the training date for both.
By Anthropic. The government blocked it after it was already released.
Every LLM product goes through testing and alignment after training, maybe even some quick improvements here and there. Kimi probably did something similar.
Put another way, if Google says they have the best model in the world but won’t release it in December I will start caring in December, not before.
if Dario and the CEO of Moonshot switched places, Mythos would have been generally available to the public 4 months ago. As a statement of fact, it wasnt, because of Project Glasswing and Dario thinking they created a superweapon that the rubes shouldnt have access to.
> First ever comment said "Further releases of Chinese models that demonstrate the gap is not growing substantially is a huge problem. The spending will be called into question."
ah yes, because i said something factually accurate and vaguely positive about Anthropic I must be a shareholder which would mean I either run a venture capital firm or am a current employee of Anthropic...
And given that Chinese models are closing the gap there are basically two thing that could be happening. One is that they are moving faster than US companies developing closed models, and two that we're starting to hit a plateau for model capabilities where all the easy gains have been plucked, and now it's not really possible to move forward at the same rate on the frontier. Of course, both things could be happening at the same time.
Chinese labs have come up with a bunch of genuine innovations: GRPO, auxiliary loss free MoE load balancing, MLA, muon optimizer, and a bunch of other ones. The Deepseek papers are really well written, this isn’t just sneaking a peek at a peer.
The problems are inherently harder now too, partially because they take longer, so your training pipeline is waiting for long completions.
Also there probably is some “distillation” (technically pseudo-labeling, which is common in ML). But I wouldn’t put too much weight on it because that was true 18 months ago as well.
That's my thinking as well. The whole distillation thing is a distraction from the actual innovation happening in this space. What will be interesting to see going forward is what types of new techniques people manage to come up with to over come the current architecture limits.
The process takes time because even when you're distilling answers, you still need to actually do reinforcement training on the model. And given that Fable and GPT 5.6 just came out there simply hasn't been much time to do that. On top of that, Kimi also does better than Fable or GPT on a lot of tasks, distillation alone can't explain that, meaning there is a difference in architecture. You can watch a talk from Kimi founder to see how Kimi was actually trained and why it performs well. https://www.youtube.com/watch?v=5CkCW1P-g88
Not to mention that US companies models constantly distill each other as Musk was forced to admit under oath. This whole narrative has just been a massive cope.
> US companies models constantly distill each other as Musk was forced to admit under oath
> This whole narrative has just been a massive cope.
So wait, US AI companies all use distillation because... it's not effective and it's all just cope? Or is distillation really powerful and they all do it, which Musk was forced to admit under oath? But when China does distillation it isn't powerful and they don't need to do it, but they do it anyway because it's fun?
Either it's powerful and everyone, including the Chinese labs, use it as a way to rapidly catch-up against the SOTA models, or it's a red herring and the huge amounts of energy spent to protect and enable distillation is all just wasted money. Which is it?
I'm saying it's a cope to claim that the only reason Chinese models are catching up is due to distillation, while pointing out that distillation itself is in no way unique to Chinese companies. I'm sorry this was too complex of an idea for you to follow.
The fun part about this is that we can see who is right in about a year. If the leading labs continue making progress at hardening their models against distillation, and then they start pulling away again, we see who is right. If China is able to pass the US and release an independently better model than anything the US has, then your theory is correct.
Both sides have extremely smart people. One side has more $$$ and exclusive access to the best chips. For progress to converge without a corresponding breakthrough suggests there's something else at work.
Indeed we will, my prediction is that we'll see a model from China that definitively surpasses any US model by the end of the year. China has an absolute population advantage here along with having a much better education system. And China now dominates in published AI papers.
The US enjoyed an early advantage due to excessive money being poured into AI which led to the current bubble, and access to the hardware that was needed to train these models initially.
At this point, neither of these factors actually matter that much. The naive approach of simply making models bigger has hit a wall, and now you need ingenuity in figuring out better architecture for them. Precisely because Chinese companies have had to deal with more limited resources, they put a lot more effort into researching different kinds of optimizing techniques. And of course, China is also catching up in chip making, and Huawei clusters are already competitive with Nvidia for training. So, that gap is closing as well.
The big difference is that an absolutely insane amount of money has been spent in the US, while China managed to do this on a fraction of the budget. The AI Investment Surge graph here puts things in perspective. https://hai.stanford.edu/news/inside-the-ai-index-12-takeawa...
My prediction is that they're going to angle to become a vendor of record for the government and get bailed out. That's the only path at this point because there won't be any competition from China in this niche.
The plateau is inevitable because their rapacious training methodologies are only viable when there are no defense in place, but information continues to evolve, which means the models will have to be continuously updated, but will be doing so with less and less freely available data.
My understanding is that the labs ran out of freely available data to train on a while ago, and now primarily rely on human data vendors such as Surge and Mercor to source their data.
Fable is still the same model, it’s still a great model, and to be honest all these articles writing and speculating on how the LLM industry is going to evolve are not that insightful nor interesting.
I don’t think one should pay much attention to them.
Those Teslas were controller-less SSDs IIRC, common in the cheap embedded world.
Modern proper SSDs for computers do their best to spread out the writes (TRIM, and other features). They still wear out of course, nothing can beat physics. But if you buy a 2TB drive, there's a lot of room to spread out the write cycles that a normal user probably never has to worry about (no normal user is going to hit 500TBs of writes that breaks a modern 2TB SSD write balancing algorithm)
But it probably should be noted that embedded controller-free SSDs have significant write amplification issues. Embedded systems assumed you were trying to save money and/or compute... and also assumed you didn't do too many writes. So they really weren't designed for the write cycles that Teslas logging system did.
The FT is usually very trustworthy, so I don't doubt the actual information. However, the question is whether Kimi K3 was only benchmaxxed or actually is SOTA-ish in real use. If the latter is the case, it may be difficult for the big US AI labs to explain their valuations.
I am especially curious whether training (or at least inference) was indepdent of NVIDIA's stack - the article doesn't say. If so, it could have widespread ramifications (geopolitical and in the markets).
Moonshot use Alibaba cloud for training and inference, specifically using NVIDIA hardware, although Alibaba also make their own Zhenwu AI chips and clusters that run on them.
Other Chinese AI companies like DeepSeek and Ziphu (Z.ai - GLM) more heavily use domestic AI chips for inferences - specifically Huawei's Ascend chips.
So, China is no longer dependent on NVIDIA, but neither has it totally cut usage of (reliance on?) NVIDIA.
I’ve been recommending the use of consistent lies about name and date of birth to online systems since Eternal September began. Very few sites and systems justify accurate PII, and even for those I often still maintain dual accounts/profiles as necessary.
That never works on Facebook though, because as soon as a ”friend” reports that ”I’m not me” then the account will be permanently banned. That also triggers for photos that’s not genuinely me, like a pet or drawing as portrait.
Unfortunately, a lot of college-age people I know are getting accounts simply for access to Marketplace, which is still unmatched compared to other local platforms for buy and sell.
Never?
Facebook is pretty overrun with what are basically fake profiles. Hell, I've been curating an alter ego on Facebook for over a decade. Built up a profile with several dozen "friends" that are all kind of interconnected and regional, but of course none of them have ever met "me" IRL, and the profile picture is a funny-ish celeb pic. Facebook has millions of legit users that are "friend collector" types, and won't think twice about engaging with an account that gently strokes their online ego with likes, "Happy Birthdays", etc.
I like using a date of birth of 1 January. It's plausible but also hopefully suspicious how many people seem to be born that day if others do the same.
The purpose of the fake birthday is not to protect random website credentials. It's to prevent someone with that data from walking into my bank and impersonating me. I started giving a fake birthday after being shocked by how little info some organizations needed to authenticate me.
If an attacker can do that, they could also do that with my real birthday had I used that. My birthday isn't a secret against anyone who wants to look hard enough. Therefore this method doesn't provide any kind of security against attackers gated only on knowing my registered birthday. I never claimed that it did.
I use the 1st of my birth month. Slightly less suspicious? It's at least a little easier to remember. Generate fake profiles and identities usually is easier when you have bits that are rooted in your actual reality. Like, you have the same zodiac sign either way in this case, so you don't have to remember two of everything. Or if you're talking about a birthday trip, or related birthday thing from the past, details about the weather would be consistent, etc.
I heard from a number of Syrian refugees that this is actually very common in countries like theirs, where births may not be recorded, records are lost or destroyed. Some people don't even know their exact date of birth and they would typically enter January 1st on forms like this too.
My first name can be shortened, and I go by either. When I first signed up to Claude, I thought we were entering the world of artificial “intelligence”, so I told it my name was “<long form> or <short form>”.
Well, it hardcodes that field rather than running it through the model, but I’ve kept it so I get an evil chuckle to myself (or perhaps pyrrhic reassurance) at its lack of smarts and a reminder that it’s still a somewhat subservient product experience that isn’t all that smart after all.
I must have made a claude.ai account when they first launched and forgot about it. Last week I logged in (through google) to get a subscription and it greeted me as "Hello, Master". I thought it was quite edgy at this day and age. :)
reply