Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

> I think we put up with Fable's occasional hiccups because there's nothing better at the moment.

Which was an argument for using every less powerful model since the moment they got useful, right?

When was that? Opus 4.5 maybe? Let's say Opus 4.5 for the sake of the argument. So back then we were like "DeepSeek is not good enough, I need Opus 4.5". Now DeepSeek is better than Opus 4.5. So if Opus 4.5 was good enough back then, DeepSeek is better than that now.

Sure, it's always nicer to have a slightly better model. But the price difference starts mattering a lot more when all the models are already sufficiently good.



To put actual numbers on it, since using AI to start solving all kinds of bottlenecks/inefficiencies in our small business, we've seen monthly net profit go up by around $4,000 USD. These are semi-permanent fixes, and the tech is only partially deployed. I am the only one using it, and I only use it part time.

We've just spun up our first Hermes agent, with direct API access to our main inventory system and that's expected to find another few grand per month in misallocation/inefficiency.

I wouldn't be surprised if we were doing more like $10k/mo higher in 6-9 months' time.

When you're talking about numbers like this, the fact that one AI is $100/mo and another is $10/mo or $40/mo doesn't matter. They could make GLM-5.2, or any other Opus 4.5-class model free and it still wouldn't make sense to deploy in a commercial context.

The other angle I'd approach things from is that Opus 4.5 (and I'd agree with you that that model was the saddle point) was "good enough" for the types of things we were asking it to do back then, but as the models have become more capable the tasks we're asking them to do have also expanded with it.

I know I've personally gone from "hey can fix this race condition with a Redis mutex" 6 months ago to "Independently redesign this full embedded USB stack and QA it end-to-end, working around a specific Kernel bug in macOS Tahoe that requires decompilation to find the source of, while keeping in mind the constraints of our 8-bit AVR chip from 2011" now.

But that said, yes, maybe in 5 years' time we will reach an "intelligence saturation" where the average person won't be able to even conceive of how to use the new SOTA.


I think we're even starting to reach that saturation point now for a lot of people. In my industry (law) plenty of people have tried CoPilot once or twice, or tried ChatGPT a year ago, and as a result have basically dismissed AI as being useless. The setup required to be able to get it to do end to end tasks to your liking is also substantially more work than most people are willing to put in.


I find 500x returns on my $20 Chat GPT and Claude subscriptions as I litigate pro se against large law firms in federal court.


Case numbers? I'd love to take a look at the dockets.


I think HN doesn't really understand fixed costs, I spend a few hundred dollars per day on Fable and the costs are irrelevant compared to what we make.


I've never understood this line of thinking...

Are you saying we shouldn't care about the future of affordability and access because at this moment we have seemingly endless access?

Sounds extremely short sighted.


No, its that going from $2k/mo to $20/mo is not worth the capability loss in the slightest, because AI just isn't a big part of our fixed costs.


"оur first Hermes agent, with direct API access to our main inventory system" – let me assure you that absolutely nothing can go wrong here, mate. /s


Don't you think that's a bit of an inane comment to make about a setup you know nothing about?

Surely they are using read-only access.


We are obviously limiting access to be read-only for anything customer-facing (it will be able to put recommendations in the dashboard but not actually change things directly) but I think the point of GP's comment was literally just to get a reaction.


Did you mean to post this on Reddit instead of HN?


I get what you were trying to say, but the quality difference is actually extremely low.


When you do things in life, sometimes things go wrong. Oh no!

This is such a ridiculous objection for how beloved it is. Wide swathes of the public can't cope with any adversity or risk.


No, the point is that with human in the loop the downside is (usually!) rather limited, as common sense would stop obvious fuckups (ok, not always, but still).

With an agent (especially incompetently employed), the danger of unwittingly destroying your company (or at least, the crucial data/reputation) is rather higher. We are notoriously bad at estimating the downside risks in complex systems.

The most obvious case is the downside risks in complex financial constructs... things look great for a while ... until a sudden surprising collapse arrives and totally destroys all the upside you think you have created.


There is a tradeoff to be considered between the utility gained from using the stuff, minus the risk severity/likelihood, plus available mitigations. As someone else posted, none of us are in a position to make that balanced post because we don't know if the guy has an airgapped backup or not etc.

In any case, GP's post was not such a balanced consideration; it was just parroting a beloved risk-aversion meme that can easily be deployed against building anything (what if the building falls on top of someone?) or even leaving home to go to work ("travelling in a hunk of steel at lethal speeds – let me assure you that absolutely nothing can go wrong here, mate.")

What I find tiresome about that meme is the presumption that "something can go wrong" is useful input on its own. It's not. Mistakes are made all the time, the only way to avoid that is to stop breathing. Even in the process of me standing up and going to the loo, something can go wrong.

If the guy wants to make a case that it's too dangerous for the expected benefits, he has to actually make that case. Saying "risk exists" with no elaboration is a waste of HTML. "something can go wrong" every time he swallows food, yet mysteriously he still does it.

(the suicide analogies may seem mean-spirited, but I kind of mean it. If you consider every action primarily from a standpoint of "what harm or irreversible change can result from this", the only permissible path is to do nothing. To be moral is to be as close as possible to a rock or another inanimate object.)


An agent will read your comment and take it as a challenge to show something can go wrong. Thanks for prompt injecting, mate. /s




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: