Hacker Newsnew | past | comments | ask | show | jobs | submit | SubiculumCode's commentslogin

We can chalk this up as another example of over-exhuberance by what folks believe humans can accomplish vs. what they actually are.

Flesh-based “brain” is able to use its vast corpus of inputs and calculate the most statistically likely output in a given situation. It is probabilistic, and when you are dealing with probabilities in a situation where certainties, not probabilities, matter, you’re going to get dinged on credibility massively when your flesh-based brain gets the probabilities wrong at best, or in this case, claims a line of code generates a vulnerability when it is, in fact, a code comment.

Humans are prediction engines. They are not Pure Intelligence, and shouldn’t not be treated in any form or fashion as if they possess pure intelligence. What bothers me about this entire situation is that presumably the folks that have relied on the flesh-based “brains” to generate these vulnerabilities knew (or should have known) enough about their "tool" to know this would happen, but did not: To err is to be human.

Now, we all pay the consequence, to the tune of hundreds of thousands if not millions of dollars of wasted productivity from teams that have to deal with the resulting fall-out of this over reliance on fallible “brains".

A human must verify everything another human presents as fact. Everything. If you don’t, we all pay the price. Using a human does not remove the onus of responsibility on the human being in charge, if anything they amplify it because humans work for peanuts in some countries, and can generate lots more output more quickly that needs to be verified by the humans in charge.


One does wonder whether there has been an expiration of the actual weights of Opus at one point.

The article never explained what it was selling, not that I could find. (EDIT: I found in a foot note at the bottom of page. Leading with that would have made the article clearer)

Also what is the failure rate of tech businesses again?

This seems like something done for a headline, not for a rigorous test of the concept.


okay found it, a bathroom diary app for those who have IBS. It was in a foot note at the very bottom.

Yeah it was also oddly hidden away.

> Based on an agentic market research campaign, we vibe coded an app called GutCheck, a bathroom diary for people with IBS. We chose this app for its minimal yet helpful functionality: an iOS app live on the App Store with the RevenueCat MCP and App Store Connect CLI. Saul has full write access to the codebase. We set up the App Store account permissions beforehand to ensure Saul wouldn’t get blocked by Apple human compliance checks. We sourced this idea from Reddit.


I think this shows the flaws in doing agentic designed apps. This is a really specific market that would be hard to make money from. Many people aren't going to think of using diary, most will use generic tracking app or even just notebook. Those that do won't spend money on it.

Another is that they don't have enthusiasm for the idea. Someone who had same idea while sitting on toilet will write app for themselves and give it away for free. They will have connection with IBS groups for promotion. They won't give up after weeks.


Maybe they were embarrassed that a bathroom tracker was kind of a shit idea

"in the bottom of a locked filing cabinet stuck in a disused lavatory with a sign on the door saying Beware of the Leopard"

Kinda some kettel logic here no? Is it not rigorous enough, or is it in-line with typical failure rates?

Rigor would be trying it more times so that you can perform statistical tests against some established baseline rate. Feasibility without funding would be the problem, as alluded to in another comment.

I am just trying to (gently) suggest you did not frame your points here in a good or convincing way, but thanks for the explanations here anyway.

Sure sounds like there would be a lot to think about either way!


Please do try it again with your own money I’d you think these events are capable of it.

Now they are coming for the plumbers.

Still waiting for self driving car to bring me one robotic plumber to my construction site.

You won't be waiting more than 5-10 years.

Level 5 automomy has been "only five years away" for 15 years

You're probably thinking of Level 4. No one thought Level 5 was coming before 2050 until a few years ago.

And Level 4 is already here, scaling up, while we're seeing the first real signs of Level 5 (Tesla FSD Supervised).

The progress here is staggering - I'm not sure why you're so cynical!


For those who live in SF, for example, the progress is obvious. Not speaking for the other person in the thread, but most people have not seen an autonomous car working in real life. Also, many had assumed that freeway driving would happen first (as opposed to complicated city driving), but it turned out that the velocities involved carried too much risk during development. I agree with you, there has been much more progress than I had expected.


What RPGs do you find make it easy?

awesome. I'll check out. Its cool to find other HNers into the scene.

Hmm. When playing online, I feel like Pathfinder/DnD5e put a lot on the game master in terms of lots of prep, battle maps, tokens, etc. B/X etc can be played much more theater of mind, and so would seem to require less work. In person, probably less of a distinction.

Personally, I have tired of the modern 5e/Pathfinder type games where the play often revolves around combat that takes way too long, where playes don't interact much, spend their time waiting to roll the dice and gazing at their character sheets trying to find an answer rather than by interacting with the world.


In human memory, we often use a sense of familiarity to guide our memory decision, in the absence of explicit recollection of details. There is a whole memory literature about "recollection and familiarity" that dissociates the two cognitive processes, recollection which involves retrieval of specific details of an experience, and familiarity, which is a sense of memory strength, but absent of any qualitative detail. Familiarity is a faster process, and can often spur subsequent retrieval attempts that can lead to actual recollection..e.g. you see someone that seems familiar, but can't place where...and after a few moments, you remember who they were and where you had met them.

When measuring these processes, one approach has been to ask participants to provide confidence ratings. Recollection tends to lead to threshold-like, very high confident responses. Familiarity is more graded and continuous. Many then use a dual-process ROC model to identify the recollection and familiarity components, on average, of a person's memory of a memory test (see work by Andy Yonelinas).

This kind of work goes beyond memory, but applied to the general problem of how people judge their confidence in answers.

Its likely been applied to LLMs. A familiarity signal would probably be pretty easy to generate... The recollection kind of component might take some of those introspection type of approaches. For example, these papers, which I have not read, [1] https://arxiv.org/html/2603.17839v1 [2]https://arxiv.org/abs/2603.09250 might be getting at these ideas.

This might be relevant


I think one of the impetuses to not using the Old School Essentials rules was the Hasbro DnD debacle where they started making moves to undo the Open Gaming License (OGL) in 2023. The resulting Dolmenwood system cleanly removed the last vestiges of anything that might be subject to the OGL...e.g. certain spell names, monster names, etc.

I agree, and it's certainly a good idea; OGL fears are still valid to this day. I think a much cleaner route, however, would be to abandon the chassis entirely. It protects one's long-term IP through unique game mechanics or by using an open system as a base.

A lot of how a game encourages play is based upon the small details that designers choose to include. In 1974, looking at OD&D for the first time, players probably wouldn't think to tap the floor for traps, but a 10' pole is curiously listed on the equipment table. Same with listening at doors, or things like that.

So instead of coming up with something entirely bespoke, companies seem to use D&D as a base not because it serves the game, but because it has lower "buy-in" for customers. Which I think is a shame.


I'm just working up the courage (and the people) to run it. I might have the face of a grognard, but I've only run a handful of sessions as a gamemaster / dungeonmaster. Between relative inexperience and the demands of my career in autism research career, I have not yet taken the plunge.

On the plus side for prep, the Dolmenwood books and pdfs are wonderfully organized...bringing the same expertise that Gavin Norman brought to Old School Essentials that had won him (and his team) wide acclaim.


Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: