Hacker Newsnew | past | comments | ask | show | jobs | submit | klooney's commentslogin

I'm not totally convinced that models are fungible, the claudes/gpts/Gemini all have pretty individual feels when you're working with them. I wouldn't be surprised if the approaches don't scale or even work the same in a poly model setup

that's anecdotal though, right? your subjective feeling of how a model responds to you will greatly influence how you 'feel' about a model and its performance in the same way that a co-worker who you get along with will fuck something up and you'll be more forgiving than when you work with a too-verbose, mansplainer of a co-worker who fucks up

I think until we have actual repeated-use measurements tracked over time (eg consistent prompts used to do the same tasks, count number of hallucinations and errors and bugs over a long period of time) you won't really have any idea of which model is better

I also think of it like a car - some just feel better to drive even if they are materially worse in other measures. until you start measuring the metrics important to you (eg MPG and cost of maintenance over a long period), you have no idea which car is actually better suited for you. and the fact that you can only do so with a limited number of cars (or hours available to work, or money to burn on tokens) means there's no true measure approaching objectivity


Anecdotal or subjective don't mean "wrong." I would 100% agree that claude and chatgpt have different 'styles.' They do have their own patterns, and those patterns are distinguishable.

it doesn't mean wrong, it's just probabilistically full of unregarded bias and prone to hallucinated ineferences, much in the same way that LLMs sometimes are

You can write a config that points your claude code at any anthropic-compatible endpoint, and replace opus and sonnet with whatever you want. Things like deepseek and glm in there do not feel at all like they do in more minimal harnesses, but neither do they convincingly act like claude models. It's very stark and frankly they run like shit this way. Try it.

it's almost like the harness was specifically designed for their model and not others :)

They said that the current models could run a vending machine profitably, just not more complex things.

A vending machine in Anthropic's office, which stands to gain from narratives like "I used Claude to generate profit".


I feel like an AI lab is less complex than a physical vending machine business.

https://www.wsj.com/tech/ai/anthropic-claude-ai-vending-mach...


Bring back wheel covers!

Is this a proprietary harness? I feel like harnesses have a huge influence on how models behave



I wonder if the inference acceleration companies will ever produce a consumer product


> But you can’t just set an agent loose in production. It can easily corrupt the database, and all it can do afterwards is apologize and promise not to do it again.

More and more, people are just gonna do it. After all, the same statement is true if employees, and they're rapidly becoming less trustworthy that frontier models.


I don't know why we're jumping to conspiracy when incompetence is right there


Perhaps we should start treating them as criminals instead of assuming that they are idiots.


Reckless and negligent idiots are criminals. They should be stopped be stopped before they get someone killed. Chernobyl happened because of massive negligence, not because someone thought a meltdown would make good PR.


That is a distinction without a practical difference.


Conspiracy and incompetence are very different.


Are the outcomes both crimes? I think that's the question to be answered.


If you consider these guys admitted they dont really have eyes on pre and post training, then incompetence really does seem more likely... especially with how fast they are moving. Its the SaaS playbook, move fast and break shit.


Laundry folding has become a doable demo for startups, and ChatGPT has been spitting out college essays for years.


> As an interviewer, I don't even know if the person I interview is a US citizen

I don't either, but my guess is almost none are


That seems like a bias you should work on.


Bring it back to Netflix


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: