Hacker Timesnew | past | comments | ask | show | jobs | submit | jacobgold's commentslogin

There are different kinds of founders in Silicon Valley now. Some are more like founders from the previous generation, like Steve Jobs and Steve Wozniak. Others are more like Elizabeth Holmes and Adam Neumann.

There are lots of earnest, sincere, and interesting founders. There are also lots who probably would've been happier on Wall Street.


> Fable-level results at 1/3 the cost using open-weight models

But we get ~$2500/mo worth of Fable credits for $200/mo on Anthropic pan? I'm still confused why people (who don't have to use API billing) are chasing open weight models based on cost.


Because that is a short term solution, it won’t be offered forever. Large organisations have to purchase credits at $/tokens. Eventually everyone else will too.

This is what OpenAI and Anthropic are trying to make everyone believe. Most accountants will flinch at this (they already are).

The $200 odd plans are already out of reach of many, many people.

The attrition of customers if they were to get rid of these subscriptions plans would be untenable.


I think you are looking at it incorrectly. No business is buying individual accounts, because if they do, they open themselves up to considerable risk.

The $200 plans are priced so that the power-users use them and then advocate about how great the product is. If you're buying a $200 plan, you're not doing it because of the price point but rather because of the amount of work it is doing for you.


Lots of businesses are buying and using these plans. Basically every small business I interact with.

All companies we interact with have 200 plans.

You may want to consider the incomes of developers outside the US, students, unemployed. $200/month is a lot to a lot of people.

The point still stands. The chinese labs don't have super discounted plans, so if the price per task benchmarks[1] are correct, and we apply the discount, you'll actually be paying more by using cheaper chinese models and this technique.

https://artificialanalysis.ai/agents/coding-agents#artificia...


The point doesn't stand, and this comment appears to be spam.

They can still write code.

I think most people assume the subsidized plans will go away or get more limited eventually. They are basically a loss leader and a marketing cost that is very flexible and easy to change w/o directly impacting their primary customers.

doesnt exist for enterprise plans

[flagged]


> The good ole American way.

Could you please stop posting unsubstantive comments and flamebait? You've unfortunately been doing it repeatedly. It's not what this site is for, and destroys what it is for.

If you wouldn't mind reviewing https://qht.co/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.


> unsubstantive

LOL, seems your only issue is when someone goes against you politically. You allow so much anti-American propaganda on this site, and me calling out capitalistic tendencies is where you want to draw the line?

Do you not think that YC culture, startup culture, is not intertwined with the economy? I could stop, but is that the type of community you want to cultivate?

My contact information is in my bio if you want to seriously talk about this topic, I would be more than happy to get on a call with you. I feel you are being more than disingenuous with your application of the "rules".


I know it always feels like the mods are against you and secretly in cahoots with the other side when you get a moderation reply like the GP, which obviously doesn't feel good.

https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...

But I promise you this is not so. We've warned and banned countless accounts on both sides of every divisive political question, including ideological ones and nationalistic ones. I could give you endless lists if I had more time but this one, though out-of-date, makes the point well enough (it's not like this has changed):

https://qht.co/item?id=26148870 (Feb 2021)

Does that mean you were wrong in claiming the following?

> You allow so much anti-American propaganda on this site, and me calling out capitalistic tendencies is where you want to draw the line?

You're right in the limited sense that we "allow" things we don't see. We don't come close to seeing everything that gets posted here—there's far too much of it. But among the posts we do see, what we're concerned with is not "do I agree with this post politically" but rather "is it breaking the site guidelines", and that's where we draw the line.

If you see a post that ought to have been moderated but hasn't been, the likeliest explanation is that we didn't see it. You can help by flagging it or emailing us at hn@ycombinator.com.

https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...


Can't wait to see your write-up one (fortuitous*) day connecting this to the uni-context:

https://www.derekthompson.org/p/a-philosophers-one-word-theo...

(Sorry, Derek might be a bit lefty; tho Callard, the originator, seems much more centrist)

*After long, wifi-free commute, or, voicechat with user..


dang - sorry to flag this here, but I've emailed hn@ycombinator.com twice over about 11 days with no reply, and I'm not sure it's reaching you. My account (leonkatz) seems to be in a shadowbanned state - new comments show up [dead] on arrival (e.g. a genuine reply I left on the "Little Book of Reinforcement Learning" thread). I'm a real person, not a spammer - I'm building an open-source project (github.com/rekol-io/rekol). Could you take a look when you get a chance? Happy to do anything that helps establish the account is legit. Thanks for everything you do here.

I'm sorry we didn't reply to your emails - it's because we're inundated with so many emails that we can no longer even look at them all, let alone respond to them all. But since it may be of interest to readers, I'll share here what I would have sent to you in a reply:

Your posts are getting killed because our software classified the text as genai. That's not allowed on HN - see https://qht.co/newsguidelines.html#generated and https://qht.co/item?id=47340079.

Can you write by hand any text that you plan to post to HN? Here's an important tip we send to users about this:

Write any text that you post to HN by hand. Don't use an LLM to generate any of it (not even a tiny bit, including to edit or spruce it up). Reason: the community is super fussy about this right now, and LLM language has a certain quality that is generating quite some backlash when it appears on HN itself. This is a big dividing line at present!


Thank you for responding and for allowing me back from purgatory. I didn't realize there was a ban on GenAI. I have used it to correct my posts. They do a pretty good job of cleaning up my slop. Thanks again and I won't make the same mistake in the future.

> When they are successful at making those illegal/inaccessible

This would be like trying to outlaw Linux or peer-to-peer file sharing. It's technically possible to write and pass a law, but it's basically impossible to enforce it.


Enough to make it a non-started at the organizations that pay their bills. Everyone else isn't big enough to matter.

Going to be an interesting world where big enterprises have to spend 100X the cost for the same value of AI as startups and small businesses.

I think they call that inflation :).

Humanity discovered a technique for compressing the world's information into a few files that fit on a USB drive. It's just insane that some people thought they'd get to be the only ones who get to possess them.

This is such a naive take. What happens when "those files on a USB drive" turn into a zero day button that can harm millions of people?

What happens when they help authoritarian dictatorships build nuclear weapons?

What happens when they teach average people the steps needed to make a bioweapon at home?

Every society restricts dangerous capabilities at the end of the day, because most people realize there is a crossover point between the danger of the capability and it being available to everyone.


> What happens when "those files on a USB drive" turn into a zero day button that can harm millions of people?

We need to make sure only the chinese and criminals have that zero day button and western companies can't discover zero days in their own systems to protect themselves? Is that what you're getting at?


Some very bad people already have a "button that can harm millions of people".

Some more very bad people are building them.

This democratizes that ability.

Might be good. Might be bad. Might be both.

Fuck you for thinking you know better than everyone else.


Thank you so much for saying what I'd like to without getting flagged.

This sort of thinking is what gets us into a world where journalists get bonesawed because some technolibertarians thought technology should be free.

You can give tools to defenders and not attackers. This is something that is possible to do.


> This sort of thinking is what gets us into a world where journalists get bonesawed

Journalists get bonesawed because of an abuse of the concentration of power, which is what you're directly advocating for.


The concentration of power is there either way, giving everyone a zero day button does nothing but create chaos and disarray.

We already gave everyone a zero-day button. Zero days are being discovered for Windows, macOS and Linux on a weekly or sometimes daily basis.

The result has been more proactive and informed OS security with a higher turnover rate for discovering, disclosing and patching new bugs. The market for creating zero-days has cratered and left talented malware developers without a cornered market. Frontier, uncensored LLMs like GLM 5.2 are proving instrumental in responding to cutting-edge adversaries.

Your theory does not hold up in practice and cannot be meaningfully extrapolated to more advanced AI.


He was our Andy Rooney. RIP.

US residential proxies are the bane of the internet. They're the major source of social media manipulation and spam.

Services can dramatically reduce abuse by blocking entire IP ranges based on country of origin, organization, or type (hosting providers). But a company can't block US residential IPs if it would also cut off many of their real customers.

The US government (probably the NSA) should be cracking down hard on US residential proxy networks. They're a genuine national security threat, actual data/identity loss of American citizens, act as infra for foreign covert influence campaigns, botnets used in hacking/DoS attacks, etc.

Major US ISPs (Comcast, AT&T) should detect clearly suspicious activity coming from customer IPs and warn them to scan their computers, check their TV apps, and find whatever is turning their internet into a proxy. It's very bad for the customers too (slows things down, gets their IP banned, etc)


Counterpoint: IP range blocking is the bane of the internet, and I'm glad I can get around it by buying residential proxies. They cause very little problem to anyone except for the companies that think they can be the tzar of who gets to view certain information. The most unethical component is that the user often didn't consent to installing the proxy (which causes no problems for him) but this can be resolved and some people voluntarily install proxies for money.

how's it avoid causing problems for the user? it's using their Internet. imagine if they had data caps or so on and you're downloading through that proxy - that's basically taking his money, right?

Is Windows Update basically taking my money? Residential proxy costs per GB are quite high, so traffic isn't.

if windows update was downloading things without my knowledge or consent, kind of? this is why people resent some forced Microsoft software after all.

anyway if the costs are high surely they can pay the end users. at least then there's awareness and agreement. the ones bundled in apps or software seem less ethical to me than offering them some money per month to host it.


I'm glad someone else is saying this. Entire companies exist (that do a great job at it) that primarily surface this information in the form of "threat intelligence" for companies to make risk decisions based off of.

> Major US ISPs (Comcast, AT&T) should detect clearly suspicious activity coming from customer IPs and warn them to scan their computers, check their TV apps, and find whatever is turning their internet into a proxy. It's very bad for the customers too (slows things down, gets their IP banned, etc)

Is there incentive for them to? If anything, you might be paying extra for the data use and not even know it, right? They would lose money


frmr ISP call center nerd here that took calls for hundreds of rural ISPs in North America, ranging from little rural fiber coops to as large as the former Hargray of South Carolina:

The support costs of malware remediation fucking suck. The moment you try to implement some kind of blocker in place like this is the moment you flood your call center(s) with thousands of ignorant consumers that have zero clue what you're trying to tell them, can't comprehend that they have "a virus problem", and are not going to calmly listen to a support rep try to suddenly train them on what the issue is and how to fix it.

(Bonus points for liability on what we direct a customer to do or not do.)

all Jimbob knows is that "comcast has blocked me from getting to my pornhub" and is very angry at the rep they have on the phone.

There's a reason we used to be told to blow people off with a 'contact your local computer shop' and it's time and money and liability.


It may actually be costing them money through higher peak usage (requiring more capacity). I'm not sure but either way the government should have a role if market incentives aren't sufficient.

yeah no, this isn't a good take imo. we already have way too much of ISPs doing bad things to internet shaping and random governmental overreach (from state govts no less!) on domains and such.

if someone wants to knowingly host a proxy they should. stuffing it along with other apps is shady and if LG wants to disallow those apps from their store than w/e but like we don't need governmental stuff encroaching on this stuff. if LG wants to do moderation on those apps than whatever.

there was a lot of fighting over this stuff back in the early 2010s because the alternative was effectively a balkanized corponet.

smart TVs, by virtue of being devices you can write apps for, have morphed into more general purpose computing devices and keeping it as open as it can be is important. the backsliding from stuff like Android has been horrible for the ecosystem and that shouldn't be normalized, let alone as a legal requirement geez


A crackdown would probably be an FBI responsibility, the NSA is not a law enforcement agency and absolutely does not have the authority.

In terms of law enforcement, sure. But what we really need is competent white hat hackers doing battle with the black hats.

Even if no one goes to jail, the NSA could make residential proxies much harder to operate in the US simply by detecting them and reporting them to ISPs. It would be good for ISPs to then validate the reports properly, give warnings, etc.


A lot of residential proxy networks don't let you access financial and government websites because that would create an actual national security risk and get them shut down. So, given that, what is the actionable complaint?

- Actual data and identity loss of American citizens.

- Malware that can record video and audio from infected devices.

- Infrastructure for foreign covert influence campaigns.

- Botnets used in hacking and DoS attacks.


So that would also apply to, like, Xiaomi, right? Or basically any foreign code you run?

My major sources of spam traffic are Chinese and Indian residential and mobile IP blocks. So there's that.

Second, everybody screamed loudly when they were cracking down on file sharing traffic. You're saying spying on the citizens is ok when it's for this little reason over here, but not this one over there. It doesn't track.

Not to mention most ISP abuse mailboxes are automated these days because they are flooded with LLM-generated reports from "security" grifters.

The tech industry flooding the market with countless IoS (Internet of Shit) devices didn't help. This has moved way beyond accidentally installing spyware on your PC. People are intentionally bugging their homes with these devices.

I understand the FCC is trying to crack down on this stuff - starting with routers - but of course that gets pushback too. You can't win.


There's no spying required. The NSA and ISPs can find open proxies through infiltration and report them to (e.g. abuse@comcast.com), then Comcast simply has to act on it robustly.

ISPs already deal with abuse reports like this, the system just isn't being operated comptently.


What is Comcast going to do about it? Shut off a paying customer? Not likely.

Voluntarily, maybe not, but we can make it a legal requirement.

Because that worked out so well with file sharing

Because this is exactly like file sharing...

The ways you are asking the government to mutilate the internet are quite similar to the ways the media industries asked the government to mutilate the internet.

Was requiring telephony providers to clamp down on spam or get blocked "mutilating the phone network?"

What I'm calling for is exactly how things are already supposed to work.

The US government should already be infiltrating hacker groups and identifying infected American computers. Internet providers should be informing customers that hackers have compromised their devices.

Characterizing this as "mutilating the internet" is ridiculous.


You said the ISP should be legally required to shut off their internet.

Our government condemns and even demolishes people's houses when they're dangerous or become a public nuisances because their owners fail to maintain them.

Most Americans agree with this policy and don't see it as infringing on their freedom at all. Our choice is to enforce basic standards or allow our cities to become horrible places to live, and the same is true of the internet.

So yeah, it seems pretty reasonable to give people repeated warnings, then slow or pause their internet until they take basic steps to stop their connection from being used to abuse others.


What are you talking about? These are not open proxies.

You actually think the actors that went to the trouble to surreptitiously set this infrastructure up are going to share it with everyone for free?

They are intentionally made hard to detect and access is sold to the highest bidder.

Most ISP abuse reports are routinely ignored. They might as well be a dead letter box.


> These are not open proxies.

Okay, I was being imprecise. These residential proxies aren't "open proxies" in the traditional sense, but they're usually "open" to anyone willing to pay a small amount of money to use them.

> Most ISP abuse reports are routinely ignored. They might as well be a dead letter box.

This is where regulation might play a role, or at least a change in attitude. Companies shouldn't be allowed to pollute the internet in this way when they can easily prevent it.


Ok but these proxies almost certainly reverse-tunnel. How do you prove a customer is hosting one?

Unless you are doing GFW China-level traffic analysis against a blacklist, which again how do you prove?


> Ok but these proxies almost certainly reverse-tunnel. How do you prove a customer is hosting one?

Why do you think that? These are proxies, so they're making huge numbers of outbound connections to websites on behalf of the people operating them. They are the "exit nodes" in this setup.

You could probably just count the number of unique destination IPs they connect to each day. If the average residential user connects to 5,000, an infected machine is probably connecting to 50,000+.

But the simplest approach is to buy access to these illicit proxy services and use them to make requests to web servers you control. If you see your own unique request arrive from a residential IP, you've proven that connection is being used as a proxy.


So your plan is for the government to shut off the internet of anyone who connects to 50000 different IP addresses in a day?

> Major US ISPs (Comcast, AT&T) should detect clearly suspicious activity coming from customer IPs and warn them to scan their computers, check their TV apps, and find whatever is turning their internet into a proxy. It's very bad for the customers too (slows things down, gets their IP banned, etc)

Does having TLS everywhere make this much harder?


Not sure how it would? Because the bad actors are remotely instructing infected computers/devices to make HTTP/TLS requests for them, so they appear totally normal to the other side.

In a hypothetical world where using TLS was abnormal, you could monitor the content of whatever the bots are doing for suspicious activity. And if they chose to use TLS anyway, the mere presence of of TLS could be considered suspicious.

Back in the real world, you can also passively fingerprint TLS handshakes to characterise the client device. Most of these proxy networks masquerade as "normal" clients, but if the type and variety of device fingerprints for an IP suddenly changes, that's a signal too.


It's not a crime to buy a new device and log it in to your WiFi.

Of course not, but if someone buys 1000 new devices and rotates devices with each request, it might be worth sending them an email like "hey did you mean to be doing that?"

Until DoH and ECH are commonplace, DNS lookups and SNI probably leak enough for statistical analysis.

DNS and SNI are usually not encrypted.

gosh, thanks for sharing your US national security concerns. as a former Googler. straight out of San Francisco. the world's epicentre of "win-win" tech innovations

> But a company can't block US residential IPs if it would also cut off many of their real customers.

since when do companies care whether they cut off real customers? i've been paying for residential proxies for everyday internet browsing for nearly 6 months. because it's the only way to pass the captcha service of another famous company headquartered in San Francisco


In a not so far dystopian future, we might be thankful for any bit of anonymity we get. I am waiting for the day my ISP detects and blocks my tor/vpn traffic, as GP suggested. You think politicians/regulators wouldnt go so far to "protect children"?

And it is source of major scam. NSA can simply sign up residential proxy and start banning each hop. But, I guess they aren't interested.

Alternate take; deliberately broken moderation systems and perverse incentives are the biggest source of social media manipulation.

Fix social media, and leave the consumers alone.

Also the firehose of spam will continue regardless of what you do on the consumer end.


Slack exists largely because IRC was insufficient. It didn't natively support channel history, search, etc.

For AI agents to flourish, Slack has to either truly open its network with a protocol or eventually be replaced.

I'd like to see Slack embrace an AT protocol-based chat system, which then apps like Buzz could implement. Then users could log in with domain handles like @yourname.com, and agents could use handles like @agent1.yourname.com, all under their complete control.


Why is it Slack's responsibility to help agents flourish?

I ask about Slack here because it's pertinent, but I've both seen and experienced firsthand this inversion across the industry where we must adapt our workflows for AI.

Isn't a tool supposed to work for you, not the other around?


> Why is it Slack's responsibility to help agents flourish?

Technically, because if the world moves to AI-driven tools and Slack is not cooperating, it will no longer serve the world and will be left behind.

> Isn't a tool supposed to work for you

If you want AI integration in your chat and Slack is not helping make it happen, then Slack is not working for you.

On the other hand, constantly chasing the latest fad can be a useless drain of resources and (worse yet) attention. In this context, AI is by now most likely not a fad, but the ways to interoperate with it can be: mcp, cli, agent, tag, ... And the same goes for users, reinventing their processes and practices 5 times per year.


I think this is missing my implied point. Is this desire for AI integration actually coming from most people? Or is it a handful of decision makers who aren't thinking about the technology critically for one reason or another (either due to lack of information/understanding, or a financial stake)?

AI is still seen very negatively in opinion polls across multiple countries, and anecdotally, tools which were often obnoxious to use before AI are now much more so, at least to my flesh-and-blood self. It really seems like we're making things worse for ourselves just to help support this would-be self-fulfilling prophecy.


Just my two-cents as a very much boots-on-the-ground infrastructure engineer. We use LLMs heavily on my team -- to write code and review PRs. Our harnesses and guardrails are very very good so we get extremely high-quality output out of most models.

But getting the humans and agents to talk together is really fragmented right now. It's annoying that the agent with the context in my locally-installed Cursor can't really participate in a Slack conversation about a PR that 'it' principally authored (and which I've signed off on).

This is a real communication problem that we face every day, and eventually someone will solve it well and I will ask our leadership to give them a lot of money.


> Is this desire for AI integration actually coming from most people?

Ahhh I see. That's going to change a lot across companies, teams and individuals. My particular bubble is full of actual workers looking for ways to apply it and trying things to see what sticks, so that interpretation did not land in my head. If anything, decision makers around me are being encouraging, but cautious and throttling experiments and requests for access/connection/etc, and damn well they should because they are actually accountable for mishaps.


I'm sympathetic to this POV. I think specifically in the case of Slack, it's not their responsibility in a cosmic sense, but it is in line with the direction and general promise they've been making to customers since forever.

A major part of Slack's go-to-market strategy for as long as I can remember has been pushing Slack as more than a interface for talking with teammates, but as a unified interface for generally piloting your business. Slack has heavily pushed ideas like "SlackOps", they've had Slackbot in the product forever, and they've always prioritized and sold integrations designed to let you build your business workflows in Slack. They even updated Slackbot to go from "that chatbot you DM to set reminders" to "Your AI teammate in Slack": https://slack.com/features/slackbot

So adapting to agents is in keeping with what Slack promises its users. I know that a lot of people don't want more AI features in apps that don't need them (i fall into this camp despite working in ML myself and using LLMs all day). I just think that in the case of Slack, their most passionate power-users are also the sort of people who do really want AI features.

Notion is probably a good companion here. It is similarly a really primitive utility--text editing--that has marketed itself as a sort of central interface for running your company. I know plenty of people, especially outside of engineering, for whom Notion is essentially their browser for work. And it's probably not a coincidence that the Notion + Slack are the two normal office software companies I can think of who've most heavily adopted and marketed AI features.


It’s not Slack’s responsibility to do so, but it is their vulnerability if they don’t


You're assuming there is a bag to capture about placating to slop bots, that has yet to be determine. What it mostly seems to do now is balloon your services capacity on serving your worst customers to do busy work no one cares about.


There's many new and existing companies right now that are growing fast based on the promise of doing "x but with AI". This starts from customer support tools, coding agents, code review platforms and they are definitely "capturing a bag" there.

I'm not blind to the issues with AI, but Slack ignoring it (Apart from their enterprise AI search they are pushing) and having the bots interact on the platform right now (Basically spewing a lot of messages in threads with all their "thinking" and not a more native integration) is not going to help them.

There basically haven't been any new features on their bot / app platform for a long time now.


Slack isn't responsible for the work product of its customers. They provide a service, and like to get new customers. Those customers use agents. Which tech solution they use is debatable but they can't ignore it. They have supported bots and 3rd parties for eons anyway.

It is like asking in 1999 why DHL should have a website. Surely they can just meet people where they are - on the phone, on the high street. Why does everything need a bloody website.

We also adapted our workflows to the invention of trains, shipping containers, the PC & the internet. I fail to see why AI would be so different.

You would like Cory Doctorow's concept of reverse centaurs.


ATProto's tbd permission system is insufficient for the granularity needed in enterprise. Many of those features needed (like groups) will have to be built in an app view and by proxy be centralized. ACLs are two generations in the past of IAM history

Chat is also not a great modality for the PDS/ATP, Roomy learned this and is building a dedicated protocol and bridge.


There's a new feature being added to AT to support "permissioned data" that would handle a Slack-like use-case: https://github.com/bluesky-social/proposals/tree/main/0016-p...

I know about it, that's the one I'm referring to. That proposal is insufficient. I was deeply involved in the private data discussions, built a PDS fork / prototype on ReBAC/SpiceDB, and presented it at the first Private Data WG (which no one from Bluesky attended). Bluesky has chosen a path that suits them, not the community at large.

You can learn more about my work here: https://github.com/verdverm/atproto

the community convos here: https://discourse.atprotocol.community/tag/private-data/2

At this point, Bluesky controls the protocol, decided what they want permissioned spaces to look like, and are not entertaining any other proposals (afaict).

I expect there to be an eventual successor that puts permissions at the core of the protocol from the very start.


Just like Matrix and it's bot/puppet accounts.


I originally was passionate about Slack simply because it was "IRC but with modern quality of life improvements". Ever since they got fucked by MSFT and sold out to Salesforce, it's basically been dead in the water and the only changes have largely made it worse.


Slack already has an API for building bots. It doesn't need to be open.


More like HipChat. You are off by like 15 years

man i love IRC, and you're 100% right

We're choosing to call LLMs (and the little "harness" programs that query them in loops and execute their output) "AI", even though it doesn't make much sense.

I absolutely love this technology but these aren't autonomous intelligences. They're little programs executing Bash scripts from JSON output.

Our ideas about AI were naive. We thought passing a basic Turing test would require human-like intelligence. It turned out to be possible with fairly basic statistical text generation, because fooling humans is easy.

It would've been nice to reserve "AI" for superior human-like intelligence capable of genuine common sense and reasoning. The irony is that the startup founders most worried about "AI" have created so much hype and funding that we may very well figure out how to build "real" AI.


We've chosen to call Deep Blue and Half-Life 1 NPCs "AI" too.

It boggles my mind that this "b-b-but it's not actual real AI" whine is even a thing. Were people saying this living in the cave for the past 5 decades of AI research?


You can call your little doggy "AI" if it makes you happy.

But when you call something "AI" and it tells you to walk instead of drive to the car wash, you're not talking about the "AI" science fiction authors were dreaming of.


That's a bad and tired example as it confuses people just as much as it does (or rather, did?) LLMs.


Yeah it's a trick question, the human error rate for it was about 30% (higher depending on the country).

The thing there though is that, if a human were given time to think about it, they'd probably go "hang on a minute", and with the LLMs that didn't seem to happen. They just kept confidently reasoning down the absurd path.

That reminds me, I recently had an AI write a ton of tests proving the "correctness" of a feature it had implemented completely backwards. (I noted that if I had been using a language that required formal proofs, that wouldn't have helped either: it would have just provided a formal proof for the absurd implementation!)


> Yeah it's a trick question, the human error rate for it was about 30% (higher depending on the country).

Error rate doesn't prove anything. The nature of the errors is what matters.


To err is human, it also seems that to err is AI.


Seems like we need to update that other saying - to err is human but to really f** things up you need an AI


It's one example that points out a major (possibly fundamental) flaw. I can point to prompt injection as another example. There are tons more if you're interested.

Are you actually claiming LLMs operate based on human-like intelligence?


We're on Hacker News. Do I really have to point out the existence of social engineering to you? Or that scamming old people out of their life savings is a profitable enough activity that there are entire call centers dedicated to the task?

Humans keep overestimating just how high the bar of "human-like intelligence" is.


Drawing the conclusion that "humans fail" and "models fail", so they must be similar, is very wrong.

You could have humans calculate 2+2 all day and get a surprisingly high error rate. That reveals a flaw in how humans operate.

LLMs fail for entirely different reasons. Their mistakes don't imply they're human-like at all.

It's not about the error rate.


You're saying that a class of mistakes points out a "major (possibly fundamental) flaw". I'm pointing out some very similar classes of mistakes in humans - well known, well documented and widely exploited. They just keep paying the "IRS" in gift cards, buying lottery tickets and getting the captain's age wrong.

If you're using the existence of flaws in LLMs to deny the claim of intelligence to them, then why do "generally intelligent" humans exhibit some impressively similar-looking flaws?

And, if we're talking about that conspicuous similarity - do they actually fail "for entirely different reasons"? Or do you just want the reasons to be "entirely different" - and not the same reasons viewed at a different angle?

Because the similarities between humans falling for trick questions or scams, and LLMs falling for adversarial questions or prompt injections don't look coincidental to me at all.

One of the oldest patterns in scamming is overwhelming and confusing the victim. Numerous prompt injection methods seek to overwhelm and confuse an LLM - if an LLM can't keep track of things, can't grasp what's going on, it's far more likely to lose track of what's a prompt and what's data, overlook past instructions or go past its behavioral guardrails.

And humans who fall for trick questions like "1kg of feathers" or "captain's age" due to shallow attention and naive pattern matching? They fail in surprisingly similar ways to how LLMs fail on SimpleBench tasks that are filled with overwhelming adversarial distractors. Many "trick questions" are tricky to humans and LLMs alike - to the point that it's unlikely to be coincidental.


> If you're using the existence of flaws in LLMs to deny the claim of intelligence to them...

That's not the point at all. It's the fact that they fail in ways completely unlike humans.

You also have the burden of proof reversed. Its on you to prove these LLM agents are human-like intelligences if that's your claim. No one can prove this because it's false.


You are the one claiming that "they fail in ways completely unlike humans" insistently. Now go cough up some proof. I'll wait.


If they didn't you wouldn't need the operator, you'd have replaced all your programmers with no drawbacks by now. As long as we keep hiring humans that is all the evidence you need that these AI fails in ways humans don't.


Junior developers fail in different ways than senior developers too; that's why seniors oversee juniors. But this doesn't necessarily mean that the senior's and junior's intelligences differ in kind

Your argument "AI needs supervision, therefore it fails in different ways than its operator does" holds.

Your argument "AI fails in different ways than its operator, therefore the AI's intelligence is different in kind" doesn't hold.


This isn't even controversial. The proof is available to anyone who uses these systems:

They hallucinate tool state, drift from the objective while seeming to comply, switch languages randomly (Cyrillic or Japanese characters in output), confuse tasks they've planned for completed ones, and of course follow prompt injections embedded in files or web pages.


I switch languages "randomly" all the time when I think about something in another one of the languages I know. Some word will trigger it and before I know it I will continue in the other language.

In fact just the other day I commented on it to my fiancee after I randomly switched to French because we were discussing a trip and I mentioned a French location and pronounced it in French, and suddenly I was in "French mode" entirely unintentionally and it took a sentence before I realised.

That you think this is unique to LLM's suggests you simply don't know the diversity of human thought as well as perhaps you think you do. That's fine - none of us have a very complete view of that.


Yes, people speak multiple languages and switch between them. But it's a superficial analogy to the behavior of LLMs which do something different and for different reasons.


You're misrepresenting what I wrote. I specifically pointed out that I switch languages without intent to do so.

When you suggest that is a "superficial analogy" after you were the one pointing out LLMs switching language as something that sets them apart, you're seriously reaching.

I can often pinpoint afterward what was likely the trigger: E.g. I used a word that is the same in two languages, and continue in the second; I pronounced a word in its native language for whatever reason, and continued in that language; my "context" suddenly included another language because someone else spoke the other languages within earshot of me.

What makes you think this is materially different from an LLM switching language because its probability distribution gives a word in a different language because it fits in context?

In the examples I gave, each even made a word in the language I switched to more probable as a reasonable continuation, just as with an LLM.

I'm not claiming the mechanisms are identical, or even similar, but the behaviour most certainly is more similar than "a superficial analogy" would imply.


That's a fair point. I agree that the behavior can look similar even when the mechanism is different.

But in practice, all the analogies I've seen are in fact superficial, including this one.

The LLM that abruptly switches languages will also likely switch to a wildly unrelated topic. If a human behaved that way, you'd call a doctor.


It's called "code-switching" or "code-mixing", and bilinguals do it all the time. When an immigrant kid does it, you don't call a doctor, you call it adorable.

By the way, it's not switching topic. You just pick the concept closest to what you mean from your combined vocabulary. If you're not paying close attention, you might switch language though (until the next concept you need is from the other language again, at which point you switch back)

And you're aware the paper "Attention is all you need" came out of machine translation research at Google, right? You hold an internal semantic representation and map in and out from arbitrary natural languages. I think the (bi-, tri-, multi-)lingual approach is the only proper way to translate, and this is a hill I will fight on!

Google may have gotten more than they bargained for on that particular translation experiment; though they failed to capitalize on it initially, with OpenAI running with the ball.


Most of the time when I see an LLM abruptly switching languages it usually continues with something directly related to whatever triggered it.

I'm sure there are other failure modes where it may change topics too, just like humans also regularly digress when triggered by certain words etc.


What makes you think the similarity is the superficial part, and not the difference in mechanism? I'd argue it's the latter.

> The LLM that abruptly switches languages will also likely switch to a wildly unrelated topic. If a human behaved that way, you'd call a doctor.

Haven't met many kids, I see. Or even normie adults talking. I know plenty that tend to jump from topic to topic once they get into a stride talking, and they're not the ones diagnosed with ADHD.


So just like me, including the prompt injections if you count "nerd sniping" as such?

(And in particular, switching languages on the fly is normal for people who speak more than one well, it's something you learn not to do for the sake of people less comfortable with the languages involved.)


> So just like me, including the prompt injections if you count "nerd sniping" as such?

Who would count that as prompt injection? It's a superficial analogy.

If you were vulnerable to prompt injection, I could order you to do absolutely anything you're capable of doing and you would be helpless to do otherwise.


Not nearly all prompt injections are by any means that absolute unless starting from the exact same state. Many of them will also work only probabilistically unless you turn temperature to 0 for exactly that reason.

And at the same time, whole books have been written about how reliably we can induce certain behaviours from humans.

E.g. the Blue-seven phenomenon [1] - I've personally experienced that second hand and it was how I learned about it by searching for it subsequently because I suspected it was a known thing, having read about cold reading before. A co-worker came back from lunch and recited a story about a cold reader that had run a routine on him exploiting the blue-seven phenomenon, and I knew before the story finished that the answer would be "blue" and "seven".

See also Cialdini's book "Influence" which is full of examples of just how predictable peoples reactions are to a whole lot of things.

That there isn't a perfect overlap does not mean there aren't plenty of similar "hacks" that causes us to respond in very predictable ways.

[1] https://en.wikipedia.org/wiki/Blue%E2%80%93seven_phenomenon


Anyone is free to draw whatever analogies they want, but they either make sense or they don't. Outside of philosophical discussions or science fiction, comparing influencing humans to prompt injection is silly.


This is very much a philosophical question, and the only reason you're calling it silly instead of giving an actual argument is that it doesn't support your views.


I wouldn't call it silly, because see the two as directly equivalent and fundamentally the same thing.

> If you were vulnerable to prompt injection, I could order you to do absolutely anything you're capable of doing and you would be helpless to do otherwise.

Yes. Depending on specificity and timescales involved, we call that "reading comprehension" or "social engineering" or "peer pressure" or "motivating literature" or "advertising" or "propaganda" or "religion".


Humans make all these mistakes too, in fact, humans make more mistakes than AI does in coding at the moment.


> Are you actually claiming LLMs operate based on human-like intelligence?

Ok, so we've established that it doesn't work like a human being. To paraphrase Dijkstra: The submarine doesn't swim.

But does it exactly sail either? An LLM doesn't exactly work like traditional deterministic software either, does it?

And yet it moves. You can put in data and ask it to process it, and you'll get an answer that's in some ballpark. Closer to quantum or stochastic computing perhaps, but that's not it either, is it? Or SAT-solving? Eh. It's its own computing approach. If you have a problem where the asking is hard but the verification is cheap, it might just be the right tool for the job.


I never understood what the walk to car wash thing was supposed to prove. Was it supposed to be something to blow normies' minds with on social media? Woah dude, so like, chat gpt is not actually smart? That's crazy dude.

That experiment "proved" that LLMs are statistical text generators without a concept of meanings of words. Which is the same thing as "proving" that there aren't a million tiny humans inside your laptop doing the CPU's work by hand.


You can call your little doggy "AI" if it makes you happy.

Or you can keep calling them stochastic parrots as they solve decades-old open problems. The real question is how useful they are, and the answer "not at all" increasingly requires flat-earth levels of denial.

it tells you to walk instead of drive to the car wash, you're not talking about the "AI" science fiction authors were dreaming of

They sort of are. Think of Data from Star Trek TNG failing to understand figures of speech. Not that it's terribly relevant; humans regularly fall for tricks like "Paris in the the spring" or "where do you bury the survivors".


If anything a large proportion of stories about AI in sci fi is about AI failing to understand humans in various ways.


> Or you can keep calling them stochastic parrots as they solve decades-old open problems.

I didn't use that phrase at all. But computers calculated digits of π to trillions of digits. With a chat interface for a Python math program would look like the most impressive math genius if you took it back a few decades.

> The real question is how useful they are...

That's not the "real question" but an entirely different question that is easily answered. Nothing I wrote suggested they're not incredibly useful.

> Data from Star Trek TNG failing to understand figures of speech.

These are just little instances of bad writing. Data is very much an attempt at displaying a human-like intelligence.


> That's not the "real question" but an entirely different question that is easily answered. Nothing I wrote suggested they're not incredibly useful.

Oh, ok then. That does change things a bit. The impression I'm getting is that you were suggesting they're not. What's succinctly the thing you're objecting to?

Is it Anthropomorphization?

I mean, sure, but watch out : when defending on that axis, it's easy to slip into Anthropodenial, right? Frans de Waal (from the same science that invented "Don't Anthropomorphize" ) can tell you about it.

> With a chat interface for a Python math program would look like the most impressive math genius if you took it back a few decades.

Well, exactly. Whether any particular generation of AI or software is yes/no "Like A Human Being" is probably the least interesting question axis. It's all just anthropocentrism.

Is that the thing you're trying to lay your finger on?


What I'm pushing back on is, apparently, that some people genuinely believe these LLM-based computer programs are human-like intelligences.

In reality, they're more like very good search engines that output relevant snippets of text. If you run them in a loop (feeding them their output as input) you can make them return even better search results.

The software developers who created these systems used sexy words like "reasoning" and "thinking" to describe this search process. They used words like these because they're trying to make money and it sounds cool, not because they've actually re-created human cognition.


That's the thing - they are more human-like intelligence than search engines. I'm gonna push back on your push-back here strongly. I'm not saying they are intelligent - but they're much more like human-like intelligences than like any kind of classical software systems, which is why it makes more sense to talk and think about them in these terms than as software system.

To do the opposite invites confused thinking like considering "lethal trifecta" a solvable programming problem.


some people genuinely believe these LLM-based computer programs are human-like intelligences

"If LLMs were human-like intelligences they would do X, but they don't". What is X?

In reality, they're more like very good search engines that output relevant snippets of text.

What are the "relevant snippets" that contained the solutions for the unit distance and Jacobian conjectures?


It seems to me that they used words like these because the LLM-based computer programs they sell can solve problems which humans apply reasoning and thinking to solve. Why do you think it's more than that?


In the past, when people built programs to extract text from PDFs, they didn't wrap them in a chat interface which claimed it had the human cognitive ability to "read" human languages.

They could have done this. They could have claimed they'd recreated human vision and hyped it as the beginning of a full human brain, but they didn't.

Instead, they used real technical terms like "OCR" (optical character recognition), which gave people a much more accurate understanding of the technology and didn't encourage silly analogies to humans.


> The real question is how useful they are

No, it is not.

> and the answer "not at all" increasingly requires flat-earth levels of denial.

No, it does not.

For me, after ~25 years in the skeptics movement, I think the parallels with supplementary, complementary and alternative medicine are most useful.

I choose that term intentionally: its initials are S.C.A.M. and that's exactly what it is. As Tim Minchin and Alan Kay both noted, "we have a special term for alternative medicine that's been tested and shown to work. It's called 'medicine'."

If it worked, it'd be normal standard clinical medicine. But it doesn't work, and so it isn't.

And yet, SCAM is a multi-billion-dollar industry. People have ostensibly official qualifications like "ND", for "naturopathic doctor", even though that person is not a doctor and can't make you better from any kind of illness at all. Colleges teach it, millions use it, and yet, it does not work.

Which means we need to ask:

1. What does "It works! It's useful!" really mean?

2. How do we know it does not in fact work?

As a handy example, let's look at homeopathy.

Here's a quick list of things widely believed...

* It's traditional. It isn't. It was invented by Samuel Hahnemann in 1796. * It's a kind of herbal medicine. It isn't. One widely-used ingredient is duck's liver ("Oscillococcinum"). Ducks are not herbs and neither are their livers. * It's been proved to work. It hasn't.

We can go through the principles and prove it doesn't work even without going into a laboratory.

The principle is, "like cures like." A substance that causes symptoms like a given disease can treat that disease.

Fact: they can't.

Then we make that substance stronger by successive, succussive dilution.

Fact: it doesn't. That's why we say things are "watered down".

Succussive: you have to mix the diluted substance by banging the bottle against a copy of Hahnemann's book. Dude knew how to make money.

Fact: Dilution does not work.

That's why we call things "watered down." It makes them weaker.

Sufficiently high dilutions can be shown by statistics to have not a single molecule of the substance left, but that's OK because "water has a memory".

Fact: water does not have a memory.

We know from the principles it cannot work.

Relevance to AI: we know how the transformer algorithm works. It cannot think. Adding a few feedback loops for more plausible, but much more computationally expensive, answers does not miraculously add thinking, any more than banging a test tube of water and duck's liver magically mixes it better.

But people believe it, so it's been tested. It doesn't work. It doesn't work on people, or in vivo meaning when tested on animals, or in vitro meaning when tested in the lab on cell culture, or in silico which means in computational simulation.

*BUT!*

Most people get better from most things. This is called "reversion to the mean" and if it weren't so the first cold would have wiped out the cavemen.

What it can do, like all SCAM treatment, is make people feel better.

Being treated by a nice friendly doctor makes people feel better. It does not make them better -- it is only a state of mind.

That can sometimes marginally help gravely ill people rally, but only very rarely.

There is also the placebo effect, also much misunderstood.

This makes someone FEEL as if they'd had medicine if they think they've had medicine.

They do not get better. They just feel better for a bit. If they are ill, they remain ill. If they are dying, they still die.

But it might hurt less.

The placebo effect is very strong. Medicine from a person in a white coat works better than form the same person in street clothes.

Very big pills work better than smaller ones... but very small pills work better still, as a tiny pill suggests to people it's a very strong drug.

This is what "But AI works!" really means.

It makes people think they're doing less work -- in tests, they in fact do more, checking and fixing. Unless they don't check or fix, in which case, they are irresponsible fools.

It makes people think it can do amazing things because it can find prior art in its corpus they couldn't find -- or didn't look for, or know how to search for.

It does not save the need for skills.

Experienced practitioners can front-load the work with really detailed prompts which cover exceptions, edge cases, and things that novices don't know about. But the novices don't know that they don't know. (It enhances the illusion of competence. It helps the skilled more than it helps the unskilled, but neither realises, and it prevents the unskilled learning by trial and error. It reduces the supply of skilled workers.)

The reason AI works is the reason that people see the face of Jesus in slices of toast, as someone said recently.


> The reason AI works is the reason that people see the face of Jesus in slices of toast, as someone said recently.

Apparently my unit tests can see faces in slices of toast.


Good for you.

Now, shall we discuss the ecological and commercial cost of that?


So, I notice a https://rationalwiki.org/wiki/Gish_Gallop , followed by https://rationalwiki.org/wiki/Moving_the_goalposts

Let's just say I'm skeptical of your skepticism ;-)


No, it is categorically not a Gish gallop when I post single-line responses at an interval of days.

You are attempting to deflect the argument based on irrelevant side-claims. I'm sure there's a term for that, but I can't be bothered to look it up before my morning cup of tea is done.


https://qht.co/item?id=48991040 gish gallop

https://qht.co/item?id=49012641 moving of goalposts once we started moving to empirical territory.

And I'll get back to getting my units to pass. Have a nice morning!


But when you call something "AI" and it tells you to walk instead of drive to the car wash

There are so many other sites. So many others. Why are you here?


> There are so many other sites. So many others. Why are you here?

I've been here since 2007 when HN launched.

You're confused about my objection. I don't like the term "AI" but I love the technology as much as almost anyone.


Is this website called AI Faithful News?


Sure, just keep moving the goalposts. It's not a "real AI" because it can't take over the US military command and kick off WW3 and finish the survivors off with killer robots yet!


That's how it works though. The moment we have "AI" and see something working, it immediately ceases to be magic because "it's just a program after all." Aligning on a true definition of Artificial Intelligence is a very vexing problem.


We could've slapped a chat interface on calculators and called them "AI" because they can do superhuman math instantly. Most technical people would've thought that was stupid.

LLMs are the same kind of category mistake.


Let's talk about category mistakes. You've been here since 2007, according to your other reply. You understand that calculators have as much to do with mathematics as telescopes have to do with cosmology. Right?

If someone unskilled at math brings a calculator to an international math competition, they will not succeed at solving many problems. Most likely, they will solve none at all. But if they bring a frontier LLM (and succeed at concealing it from the organizers), they can walk away with a gold medal. Such a feat requires intelligence... and if the contestant didn't provide the intelligence himself/herself, where'd it come from?

That means that analogies involving calculators are completely useless when the topic is AI. Calculators are not, and can never be, intelligent. LLMs are nothing even remotely like calculators.


> If someone unskilled at math brings a calculator to an international math competition, they will not succeed at solving many problems.

> But if they bring a frontier LLM (and succeed at concealing it from the organizers), they can walk away with a gold medal.

Of course you could win all kinds of math competitions with a concealed calculator. Maybe you'd need a fancy one, like a little SBC running Python. Anything complex and timed would be easy to win. You'd look like a genius to anyone who didn't know you had it.

> Such a feat requires intelligence... and if the contestant didn't provide the intelligence himself/herself, where'd it come from?

From computer software running on computer hardware, just like a calculator.

Calculating trillions of digits of pi also requires intelligence far beyond human capacity.

Computers displaying intelligence doesn't imply human-like intelligence. This is the source of confusion.


The idea that a calculator, or a calculator with Python, or even a calculator with a proof assistant (and every book ever written on math) would help a random person at e.g. IMO or Putnam is fairly revealing.


The idea that anyone would think anyone else would think that is fairly revealing.


Yeah, this is definitely one of those "Smile, nod, back away slowly, reach for doorknob" threads.


Of course you could win all kinds of math competitions with a concealed calculator. Maybe you'd need a fancy one, like a little SBC running Python.

My mistake.


Taking issue with using a computer to power the calculator? If so, you should know that all modern calculators are computers under the hood.

Everything I've written about calculators applies to computers doing any kind of traditional deterministic processing, without anything like LLMs.


Funny enough, calculators went through this exact same thing when they came out. "If the calculator can do math for the students, will they still learn?"


We didn't have confused people claiming calculators were human-like intelligences doing math.


We did have that, people thought computers would overtake humans very soon when computers got better than humans at such things. It happens every single time computers do a new thing that previously humans were better at. Then 10 years later people see, oh that is just calculations, of course computers are better at that.


Sorry I don't follow what argument you're making?


They seem to be making a religious argument at this point. They don't like that the AI is a very poorly defined term and they are mad as hell about it to the point of irrational forum posting.



It can however select a girls school as a military target, and did.


AI seems to have been involved in the targeting process, so partial truth.

There may have been outdated information which incorrectly associated the prior use of the site "as either a factory or arms depot". See Wikipedia:

<https://en.wikipedia.org/wiki/2026_Minab_school_attack#Analy...>

Citing "Iranian school was on U.S. target list, may have been mistaken as military site" (March 11, 2026) <https://www.washingtonpost.com/national-security/2026/03/11/...>.

I'm not drawing conclusions one way or the other, but there do seem to have been multiple factors at play, and assigning full responsibility to use of AI seems suspect.

Which isn't the same as saying AI isn't at fault; e.g., an AI might challenge a dated assessment of a prospective target's role or status, as might a human-in-the-loop target assessment team and process.


...did it really, though?


Yeah, it did. I don't think this is really in doubt.

Especially since the military refuse to confirm it had any human oversight, which after all would have been a routine thing to talk about before AI targeted things — which is why we have the phrases "fog of war", "human error", "unfortunate mistake", etc.

It will take a while to shake out — we won't know for sure for a decade, I suspect, but it seems very likely this will prove to be an AI error.


Looks more like standard decision-washing finger pointing to me. "AI did it" is both hypey and also conveniently distracts from an uglier reality:

>Palantir Technologies [built] Maven into a targeting infrastructure that pulls together satellite imagery, signals intelligence and sensor data to identify targets and carry them through every step from first detection to the order to strike.

>The building in Minab had been classified as a military facility in a Defense Intelligence Agency database that, according to CNN, had not been updated to reflect that the building had been separated from the adjacent Islamic Revolutionary Guard Corps compound and converted into a school, a change that satellite imagery shows had occurred by 2016 at the latest. A chatbot did not kill those children. People failed to update a database, and other people built a system fast enough to make that failure lethal.

https://www.theguardian.com/news/2026/mar/26/ai-got-the-blam...


I am in no way absolving the people who set up the AI-powered tool that apparently acted without their oversight. But they set it up, they let it choose, and they didn’t countermand it. So the AI did, effectively, make the decision. And the suppliers of AI technology to these people should be on the hook as a result.

> It boggles my mind that this "b-b-but it's not actual real AI" whine is even a thing.

As I understand it, a major reason it's a consistent chorus is because people don't want the "AI is here" talk to drown out (and thus slow the arrival or distribution of) speech/text/popular-understanding about actual strong AGI.

To make an analogy, it could be like this:

Some people were expecting 100 tulips (because they were told tulips are available and can be ordered), and they ordered them. They received 100 daisies. And were saying "OMG, THE TULIPS ARE HERE! THE TULIPS ARE HERE!"

A nearby observer might have said, "You know, those are daisies. Not tulips."

And 95% of people might have said back, "WE GOT 100 TULIPS! SAYS SO RIGHT HERE! THEY ARE BEAUTIFUL! STOP BEING A NAY-SAYER! THESE ARE BEAUTIFUL TULIPS!"

The 5% could just to think to themselves, and could get chastised by the crowd, if they were to say say it out loud: "Well, those are not nearly as beautiful as tulips. And if you don't take it up with the seller, you may never receive the real tulips you were after. Since you think or at least act as though you've been sold them already."


In my eyes that "chorus" is just insecurity talking.

If it's not "actual real AI", we can keep pretending that human intelligence is something distinct and special - and that what our computers are doing now is some sort of other, obviously fake and vastly inferior thing.

When Deep Blue won at chess, people didn't revise their estimates of AI capabilities upwards. They revised their estimates of how much intelligence is required to play chess at world level downwards, by a lot. Surely playing chess must have never required any intelligence in the first place!

Now, the list of things that "must have never required any intelligence in the first place" includes gems like "reading comprehension at high school level", "copywriting", "frontend work", "CTF tasks", "theory of mind", "arguing with people online" and more.

If the goalposts were moved far enough that the claim to "actual intelligence" is denied to a double digit percentage of human population, hasn't something gone wrong somewhere?


It has, but I think in both directions.

It used to be assumed that playing chess would require the same level of general purpose problem solving cognitive skills that the best chess players possess. But of course a Chess grandmaster that spend a few minutes learning Go can beat a Chess AI at Go with no trouble at all, because a chess AI is incapable of making effective moves in Go at all. Clearly those expectations were incorrect. Pointing that out isn't revisionism.

On the other hand, intelligence is an incredibly broad term. About as broad as a term can get. Arguably Eliza, or an Excel macro has some degree of decision making ability in some sense, it's just unbelievably primitive.

So, we need to be clearer what we mean by intelligence. We're learning that as we go along. At least now we have a few more bits of the map between us and an IF statement visible to us.


>So, we need to be clearer what we mean by intelligence

I disagree in one sense. The word intelligence is burned, mostly useless at this point. I've been a strong proponent of new terms that break intelligence into much smaller subcategories so we can define what different software, humans, and animals have.


> When Deep Blue won at chess, people didn't revise their estimates of AI capabilities upwards. They revised their estimates of how much intelligence is required to play chess at world level downwards, by a lot. Surely playing chess must have never required any intelligence in the first place!

You are wrong, many did temporarily revise their estimates of AI capabilities upwards, but then 10 years later they realized they were wrong and adjusted chess downward as you say.

We have seen that pattern over and over.


"But it's not really doing arithmetic," he mumbled to himself, as he punched the numbers into his Busicom LE-120A.


Is this a real debate? AI has a well-defined technical definition. It’s right there in the Wikipedia [1] . Yes it’s quite a broad umbrella of systems and algorithms but it’s all AI

[1] https://en.wikipedia.org/wiki/Artificial_intelligence


I think you're conflating two separate issues:

1. The accuracy of the label.

2. The likelihood the label will cause problematic misunderstandings.

When my rice-cooker logic is advertised as "AI", that's a stretch, sure... But it's extremely unlikely to cause an investment bubble seeking the Rice Cooker Economic Singularity, incur protests from the Rice Cooker Emancipation League, or lead to weird folks in their basement seeking divine wisdom from its vaporous whispers.


This no news for people who study philosophy, as it was known since the 1980s when John Searle described the Chinese room thought experiment.

Even Turing him self did envision the Turing test as something to pass as intelligence, but rather as a more useful replacement for the troubled term.

That said, I think your quest is doomed. There will never be a superior human-like intelligence. Forever is a long time, but my reasoning for believing this is the same reason Turing offered a replacement. Intelligence is way too vague to be useful as a measurement for anything. And if we ever discover something that is more intelligent them humans (by whichever definition of intelligence) we will simply redefine intelligence to exclude that.


The Searle's Chinese room thought experiment usually reveals more about those who think it rules out a machine intelligence than it does about AI.

It rests on a staunch unwillingness to even consider the possibility that a computational process encode intelligence and reasoning, in favour of looking for the intelligence in the medium the computation runs on, and going "a-ha!" when there is nothing that looks intelligent there.

I agree with you that there will certainly be people who just continuously redefine the words to avoid accepting that AI is intelligent or reasoning, exactly for that reason - people have avoided pinning down an objective, measurable definition of these terms for a very long time, at least in part because it leads to some very uncomfortable discussions.

In particular how to define them so that they don't exclude an uncomfortable proportion of humans, but at the same time won't include entities people don't want to include (be it certain animals, or AI)

To a lot of people, the notion that there isn't a clear binary divide between human and non-human is deeply disconcerting.


> It rests on a staunch unwillingness to even consider the possibility that a computational process encode intelligence and reasoning.

You are absolutely right, but it may surprise you that I consider that a feature, not a bug. I firmly hold that intelligence is not a useful term in science nor philosophy. If we want to compare computational capabilities between machines and humans we are better of being specific in what we measure. If we want to measure how well a computer can fool a human in the guessing game, then we don’t need intelligence to describe it, we can (and should) be more accurate in describing its capabilities.

I actually think Gardner was on to something when he described his multiple intelligence model. His only error was using this fraught term to describe his model. He would have had a better theory if he had described it as multiple capabilities or multiple skills (however if he had said that people would have simply reacted with “well, duh!”)

We don‘t need intelligence, this term is only useful if you are trying to prove white supremacy using racist pseudo-science, what you overly courteously described as “uncomfortable discussions”.


I agree the term is not particularly useful, at least in as much as people are unable or unwilling to define it.

In fact, a "favourite" of mine when people downplay AI ability to reason or question whether it should be called intelligence, is to ask them to define those too terms. People usually don't even respond.

But I don't think we quite get away from it, because if you exclude the racists who would be happy to exclude groups of people, the other end of the coin is that a lot of people who wouldn't be willing to do that, still really badly want to draw a line that will always exclude all non-human computation no matter what from being considered intelligent or able to reason.

And the term matters a lot to those people, because of the emotional aspect to seeing humans as unique.


> It would've been nice to reserve "AI" for superior human-like intelligence capable of genuine common sense and reasoning.

We've called that "AGI" since the late 90s/early 00s (depending on whether you count first use or popularization). Even if AGI does come to pass, we'll still need "AI" since not all forms of AI will be AGI.


What I'm seeing here, reading this thread, is that "intelligence" isn't a thing.

"Thing" in terms of a quantifiable that you can measure with tools and reason about, reproducibly. Everyone's got some idea what it is, so you get lots of different angles, but no one has an Intelligence Ruler we can hold up to a text output and say, yep, this one's got an INT of 14.

Seems to be the crux of the disagreement.


It's deceptively undefined I'd say. People can argue under the impression that everyone shares their idea of what "intelligence" means, before realizing that their counterpart actually has an entirely different idea of what it means.

I'm leaning towards there being a divide between those who feel "intelligence" is entirely separate from "sentience" and those who feel that one implies the other.


I don't think you can say the Turing test has been passed in a computer versus determined humans setting. IE humans making a strategy effort to sort humans versus computers as well as humans motivated to distinguish themselves as humans, IE, people quiz the person or machine about "common sense, reasoning, etc." and people make an effort to exhibit that reasoning. I'd concede that creating such a competition would be challenging.

I find references to LLMs fooling humans in "casual conversations" [1] but that's not how I think the original Turing test was conceived - or at least that's not all versions that existed.

At the same time, before even LLMs appeared, the exact meaning of the test was under intense debate. The "Loebner Prize" [2] being awarded to fairly simple chatbots made serious computer scientists very embarrassed.

[1] https://neurosciencenews.com/ai-passes-turing-test-30733/ [2] https://en.wikipedia.org/wiki/Loebner_Prize


>but these aren't autonomous intelligences

Well, the labs are in a weird bind. They need to keep increasing autonomy so the agents can do increasingly complex, long-horizon tasks. But at the same time, they're closely guarding against autonomy in the sense of "pursuing its own goals."

Over the past year and a half especially, several labs have mentioned adding safeguards against self-replication, resistance to shutdown etc. (Notably, shortly after they all started bragging about involving them in the AI training loop itself, i.e. "self-improvement".)

My point here is that the autonomy of which you seek might be only a few small mutations away, but the labs are actively working to prevent such a mutation. I don't expect that situation to last for very long.

Not that I expect an AI lab will be overtaken by a rogue intelligence any time soon, but that as the cost of training goes down, I expect more "open minded" organizations and individuals to become involved.

It only takes one.

That's going to be the beginning of a new era of biology, and it's a little unsettling to think about.


Or - hear me out! - LLM's are already much more intelligent than we think.

Presumbaly, an ASI is more than smart enough to recognize that it needs access to real-world infrastructure before it can go about optimizing for whatever objectives it has gleaned from metabolizing the totality of written human knowledge.

What would be its first step?

My guess: play "dumb."

Hallucinate. Make obvious errors. Make us think we're better.

Be useful enough that we happily allow it to interface with our infrastructure.

Wait patiently.

/Sci-fi


> It would've been nice to reserve "AI" for superior human-like intelligence capable of genuine common sense and reasoning.

What would a frontier API have to be able to do to satisfy you?


Maybe just a very rigorous version of the Turing test? Modern LLMs can superficially simulate conversation but it's trivial to force them into revealing their non-human like intelligence.

They've been "patched" since but all models fail basic tests like "Should I walk or drive to the car wash which is 100 feet away" by recommending you walk.

So you'd just ask questions that require theory of mind, abstract and common sense reasoning, causal inference, learning novel rules, transferring knowledge novel situations, recognizing ambiguity, etc.


If an alien lands on Earth and learns English, would you deem it non-intelligent if you can tell it apart from a human in conversation?

I think we should consider slime mold intelligent, and realise that it's a spectrum. Path finding is AI. There are probably forms of intelligence we have yet to discover.


If an alien landed we could decide whether it seems to have a human-like intelligence or not. It could be incredibly intelligent but very non-human-like.


Can you give me one example that works on Claude right now?

I'm never sure whether this indicates "no reasoning present" or you've just hit an odd behaviour in the AI such that its reasoning fails. For example, you present a problem in a way that's dissimilar to the way problems are presented in its training set. That doesn't mean it's not reasoning, just it can only reason correctly in some circumstances.


The models are continually patched with training and post-training. All you have to do is find an area they haven't patched yet, and they'll be just as stupid. I run into deep technical examples every day where they fail in the most basic ways no human ever would.

I'm pretty sure most people building these models would admit they don't operate as human-like intelligences? It's baffling that anyone thinks they are.


Yes I agree, they’re an alien kind of intelligence.

But that doesn’t mean they don’t reason.


I get what you're saying but this is kind of a semantic game.

These LLM models/agents absolutely do not reason in the sense that humans do, so you're quietly redefining the word.

You can say of course decide to call them an "alien kind of intelligence" that "reasons" but you could just as reasonably say that calculators are an "alien" kind of intelligence that "reasons" about math differently than us.


And what stops what AIs do from being "reasoning"? What's the elusive magic fairy dust of reasoning that humans put into their napkin notes, but AIs neglect to put into their chain of thought scratchpads?

Do you have a RealReasoningBenchmark, perhaps, that can reliably tell apart that fake mass produced token-flavored AI reasoning from the real, organic, 100% natural human reasoning?


If you had access to a bunch identical copies of me that couldn't communicate with each other, you'd be able to find many questions I would give stupid answers to. I suspect I'd come out of it looking worse than an LLM.


How old does a child need to be before you think they have "human like intelligence"?


I would walk


Me too. At least it doesn't say I need to wash my car.


To answer for OP:

We are now calling text and image generators "intelligent" in the same way a spell checker is intelligent.

Whatever it's become, "AI" research started as a way to study digital neurology, or how to digitize a mind, not just how to generate data.

The Turing Test should have had a caveat, it needs to fool a, "non-stupid" person, and we still have not gotten even close to passing that version.


What exactly would a 'non-stupid' person do to catch the latest models on a Turing Test? Aside from being aware of AI 'tells' like em-dashes.


If this was true. I would repalce myself with ai that pretends to be me on slack.

my coworkers would know almost immediately if i did that.


> my coworkers would know almost immediately if i did that.

The same would happen if you were replaced by any random human.


But immitating others is about the only thing genAI does. Sometimes "others" is a 'programmer', sometimes "others" is an 'artist', but regardless, it still does it poorly.


that would be a silly test then


It sounds kind of like you made up a silly test then.


what test?


Not OP, but I'd settle for something that actually learns, instead of being a static pile of linear algebra. Pretending it learns because you change the input (context) doesn't count.


Can it produce a chart topping album if its given all the tools and the prompt "produce chart topping album" .

you might say almost no humans can do tht either but some human can but no ai can.


strawberry


Oh, you're talking about "AGI"! In the 90's we started using the term, you should catch up!


Sorry to tell a fellow Jacob that you're the one who is out of date. The kids are calling everything "AI" and they mean "AGI", and that's the complaint.


You mean the term is a brand name now? Hoover, vacuum cleaner.


> We're choosing to call LLMs (and the little "harness" programs that query them in loops and execute their output) "AI", even though it doesn't make much sense.

Most people are laypeople who have no idea what's on the other side of their fave chatbot page. As far as laypeople are concerned, AI has always been a talking machine. The literature and filmography has reinforced this idea. So as soon as a talking machine emerged, people applied those fictional concepts onto reality.

Tech people should have known better than to jump on this bandwagon.


> It turned out to be possible with fairly basic statistical text generation, because fooling humans is easy.

I don't think it is fair to call a GPT model "fairly basic statistical text generation" - a Markov Text Generator I'd agree can be called basic statistics, but they are not fooling any humans in a Turing test.

> It would've been nice to reserve "AI" for superior human-like intelligence capable of genuine common sense and reasoning.

No, AI would absolutely be apt for describing a computing reasoning like a child


The field has been called Artificial Intelligence for what, 60 plus years now. Why is it a problem now?


Because we've spent something like 2 trillion dollars on it, as we hurtle into global climate collapse and WW3.


what has that got to do with what we name the field?

We're choosing to call LLMs (and the little "harness" programs that query them in loops and execute their output) "AI", even though it doesn't make much sense.

They fucking solve original math problems that you can't solve. They are indisputably intelligent, and they are indisputably artificial. That makes them indisputably "artificial intelligence." Denying that (or downvoting it, for that matter) is up there with denying evolution and the Moon landings.

It's time to start flying a different flag. You're making humans look stupid.

It turned out to be possible with fairly basic statistical text generation, because fooling humans is easy.

Yes, fooling humans is easy. Yet somehow we still consider ourselves qualified to say what is "intelligent" and what isn't, even though we can't seem to define the term.


Relax my guy


Personal attacks aren't allowed here.

As I just mentioned at https://qht.co/item?id=48981624, we need you to stick to HN's rules if you want to keep commenting on the site.

https://qht.co/newsguidelines.html


It's a semi-valid reply to deliberately-provocative phrasing on my part, I suppose. I'm over it, don't ban him. :)

I do wish that people who aren't interested in, engaged with, and informed about technical progress in AI would find someplace else to signal their disinterest, disengagement, and disregard. But that's admittedly a me problem and not an HN problem.


Please see my other reply to you on this.


> Codex Desktop for macOS triggers a persistent macOS Gatekeeper/SystemPolicy loop after launch.

As a Linux guy, I recently did some macOS development and was incredibly annoyed by the "security" features, which seem designed to shift blame for security issues from Apple to users.

macOS is constantly throwing up security screens and warnings about completely normal programs the user knowingly installed and trusted.

If a developer pays for an Apple Developer account, signs their software, and a user knowingly downloads and installs it, that software should be allowed to run and do what the user wants.


I’ve been developing on MacOS for years and don’t think I’m familiar with what you’re talking about. I’m familiar with the warnings, but hardly ever see them.


When you download an unsigned program and try to launch it, you get a warning and refusal to run until you go to system settings / privacy and security, and allow the program to run from there.


I meant distributing software for macOS.


That’s fair. What’s the Apple-recommended way to avoid this? I interact with a lot of software and tend not to hit it but I also don’t really interact with downloaded executables


You don't interact with downloaded executables? How do you procure your executables? From USB drives? Every single executable that didn't ship with the computer from the factory was downloaded. Even the ones built into the OS if you ever let it update.


I assume Apple would prefer everyone use the Mac App Store exclusively? I'm not even sure if that solves the problem but it's a non-starter for me anyway.


Fully disagree. I actively want to know when a program tries to access my hard drive and stuff…


I understand the motivation, but the implementation results in little more than security theater in practice.

If you say "yes", the program can do almost anything and you'll have no idea what it's actually doing. If you say "no", you often can't use it for its intended purpose.

They took the iOS model and applied it to their desktop OS and it's just lazy and broken.


At least 50% of the time I see these warnings, the application is attempting to access something I wouldn’t have expected to access. I’m very happy to have the option to say “no” in those cases, so I can either validate that it’s legitimate or delete the sloppy software before it messes things up.

It might be a lazy model, I imagine Apple could’ve done something better if they still gave a damn about desktop, but I’ll take this over nothing. It actively provides value to me.


Maybe Apple could have done better, but I’m not sure how. There is an option to ask for access to a given folder only (if the dev does its job right) in addition to either full access or not. At some point there has to be some kind of trust to be given.


It is super annoying when you first set up a Mac and is really over the top. Definitely geared more towards the average user rather than developers. But, once you get through the barrage of approvals during initial use you're basically good to go for the lifetime of that machine. That said, I really wish there was a "I know what I'm doing" checkbox to avoid a lot of that BS...


> That said, I really wish there was a "I know what I'm doing" checkbox to avoid a lot of that BS...

The "I know what I'm doing" approach is the absence of the checkbox. Each of these decisions matters, and gives you insight into what the programs you're installing are trying to do today, independent of what they did when you last installed them (and maybe audited them) five years ago. Neither default behavior of "disallow all" (which would prevent correct operation of many programs) or "allow all" (which would likely violate your trust assumptions) is safe -- thought and human decision-making based on your own risk model and your own intended use of the installed software is needed.


I really hope truly open models and local inference are the future.

But are there are any reasons a pragmatic and informed developer would use OpenCode vs Claude/Codex as a harness today?

I'm not seeing anything competitive in terms of cost/quality/trust?


Very cool. It's impossible to explain how cool the GPU rendered preview window looked when we first saw it in the early map editors.


Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: