FSST is based on a fixed size (255 items) dictionary of high frequency variable length strings/substrings (learned from the corpus) encoded as one byte.
Per Xiaomi, MiMo v2.6 training run cost $3.47m. A far cry from the estimated costs ($100m+) for the Big 5 (MSL, xAI, GDM, OAI, Ant). I wouldn't be surprised if salaries and R&D costs have similar drastic disparities.
For a model that matches Muse Spark 1.3 in benchmarks, MiMo v2.6 Pro is incredibly cheap, given its cache rates will remain $0.0036 per million.
I sorta got the impression that the $3.47 million only covered post-training , given that few of the graphs start at zero. Is a barely-trained model going to score 48 on DeepSWE v1.1 ?
Kano predicted that users' perceptions of satisfaction with a feature will shift from delight to expectation over time. This is either because they have got used to it or because competitors have started to include it in their offerings. In the case of the touchscreen, for instance, the newness that came with the iPhone is no longer a novelty; the touchscreen has become the norm. It's no big deal anymore, but taking it away would be a big deal!
> got lots of clever behind the scene tricks like Google or Deepseek writeups) and benchmark scores
Xiaomi MiMo is led by Luo Fuli, a former Alibaba & DeepSeek employee. Perhaps it is due to Luo just how similar Xiaomi's tech & GTM approach is to DeepSeek's.
My son, Jeffrey, was a very tiny, very sick premature baby, born Feb. 9, 1985, at a gestational age of 25-26 weeks. During the almost two months of his life, he was on a respirator, with several lung diseases, a heart problem, kidney problems, and a brain bleed.
... Jeffrey had holes cut on both sides of his neck, another hole cut in his right chest, an incision from his breastbone around to his backbone, his ribs pried apart, and an extra artery near his heart tied off. This was topped off with another hole cut in his left side for a chest tube. The operation lasted 12 hours. Jeffrey was awake through it all.
The anesthesiologist paralyzed him with Pavulon, a curare drug that left him unable to move, but totally conscious. When I questioned the anesthesiologist ... she said Jeffrey was too sick to tolerate powerful anesthetics. Anyway, she said, it had never been demonstrated to her that premature babies feel pain. She seemed sincerely puzzled as to why I was concerned. It turns out that such care, or lack thereof, is possible because, as a neonatologist explained, babies, unlike adults, don't go into shock no matter how much agony they suffer.
"such care, or lack thereof" is heavily loaded; about the only objectively-sounding statements in the last paragraph are that Jeffrey wasn't able to tolerate anesthetics, and that "babies, unlike adults, don't go into shock", both of which add up to possibility of care where alternative was to leave the child to die.
I have used Kimi 2.5 and GLM 5.3 (& 5.3 Flash). Do not need them for what I do outside of spec hardening (basically, a lot of chatting).
I tend to know exactly what I want and most of the weaker models are enough to get me there. I have mainly been using MiMo, DeepSeek V4 Flash and MuseSpark Contributor over the last month or so.
FSST is based on a fixed size (255 items) dictionary of high frequency variable length strings/substrings (learned from the corpus) encoded as one byte.
reply