Hacker Timesnew | past | comments | ask | show | jobs | submit | deepsquirrelnet's commentslogin

If you look at papers on benchmarks, they're usually created to expose gaps in how models are trained. It should be no surprise that models get better on them over time, because you can't get better at what you don't measure.

Cherry picking the benchmarks you present is where the falsehoods lie.


Praise FSM. Without circular investments we'd all be broke.

China is doing open models better. I think there's little going in Hagueseth's head other than the usual reactionary nationalism reflex.

I think it's hilarious. LinkedIn is rushing to de-legitimize themselves so hard that they're inventing a new market for someone else to step into. Apparently indeed doesn't want to take it... not sure what's going on there.


> You've used 92% of your Fable 5 limit · resets Jul 12, 12pm

So generous.


Turns out it's easier to make conspiracies than effective policy. Who knew?


The moment MTG was elected and I realized that the “true believers” were getting into Congress I knew it was going to get rough.


Competence and conscientiousness are the ideal traits but I think I’d tend to prefer “true believers” over “inauthentic charlatans”.


I think this is only accurate when no external ideas are used, but I'd like to suggest that nearly all new discovery is built on a combination of old ideas and LLMs are really good at the latter.

If you bring something new to the table, then in my experience, AIs are really good at helping you ground it old ideas. If you want to set it and forget it, then you will get the mean. If you want to do something new, in my experience, they are enablers and not blockers.


Can anybody find trustworthy stats that these actually reduce crime? All I see are occasional anecdotes about how they were used to find one person one time.

Skeptical me seriously doubts this is an effective solution for crime. But maybe that's because this country has a history of being willing to do a million expensive and privacy violating things, and only if it's a punitive measure.


I'm not sure about reducing crime but most American police departments have difficulty finding staff. Generally boring job in most places and not really liked in other places (status loss). Speeding things up is one of the ways to deal with it.


> Speeding things up is one of the ways to deal with it.

Making it a well-paid, high status job is another way to deal with it.

Not easy. Not cheap. Involves fixing quite a few incentive structures as well and weeding out corruption... Yeah, I guess you're right, speeding things up a bit at the cost of everyone's privacy and liberties is going to be what they go for.


Yeah I realised the same at some point, most people don't care anyway and we don't really have privacy regardless.

There's no salary you can pay that attracts smart enough people to these jobs in some places (while being fiscally somewhat responsible). It's similar to the problem with doctors in rural areas where wages don't matter.


UK cctv and China's system are probably the closest examples?


I don't have stats, but most police have made it pretty clear that they're used for investigations that would otherwise have very little to go on.

I don't think anyone other than the manufacturers have made claims of cameras reducing crime. You can put all the AI bells and whistles on them, but they're still just cameras.

They're a fallback option, not a dragnet. The police are generally reactive to reports of crime, not proactively trying to piece together the details of everyone's lives and nail them the moment their dog poops on the sidewalk. No AI can even do that anyway and it would be a waste of money.

There are two vocal camps of people on these threads that are eroding HN: fearmongerers and grifters. I don't understand how it got this bad, but that's the real crisis here.


You are absolutely correct but you won't get anywhere here.

I have relatives who are cops and lawyers and city councilmen. No cop is sitting in a back room somewhere tracking all the cars on every street trying to do, uh, whatever it is people here are claiming they are going to do to them.


An obsessive stalker police officer that's angry at their ex girlfriend moving on and finding a new partner will probably not be interested in EVERY car, just the one she drives and the one her new boyfriend drives.

I won't speculate as to what your law enforcement family members may or may not be capable of when it comes to this technology, but I will speculate on what they will likely do if they found out about an obsessive stalker police officer that's watching their ex-girlfriend and her new partner using this tech: they will likely assist in hiding it so as to ensure that the optics of the justice system are not marred.

The reason I suspect they will behave in this way is not because they're bad people - but because they're likely normal people who are subject to normal influences and incentives. There will be no personal benefit, and significant personal risk associated with whistleblowing on this hypothetical officer, and so they will find rationalizations for why they shouldn't. Why it's fine to let this "one bad apple" go for the greater good of the optics of the justice system.

So it goes.


I also wonder what makes people think the cops are going to trust AI any more than anyone else. A mistake on bad information is even more dangerous for them and often makes national news.


Judging by the number of news stories in which they have done just that, they will happily trust AI. A certain relatively small percentage will generate national news and that kind of blowback, but by and large, they have the guns, the SWAT teams, and the local prosecutors, and the consequences are minimal.


How do you know what a "relative small percentage" is if there are alleged unreported incidents? What would even allow anything to go unreported vs just ignored (because nobody cares)? Why would having a large capability to respond with force be relevant to whether the cops are going to listen to the slop machine?

I'm trying to understand what makes you so sure of your opinion. Unless you're living in a third-world country, from the perspective of anyone who has barely even known a cop, this sounds wildly out of touch. There already is a ton scrutiny on everything they do. Picking apart the long tail of debatable outcomes you don't like is not evidence of corruption. Corruption would be no debate at all.


Except for of course, the more than a dozen known cases where cops have been using Flock to stalk people [0]. Realistically, it's very likely that most of these cases do not become known.

[0] https://www.404media.co/cops-keep-getting-arrested-for-using...


If they don't reduce crimes what do they do? Oh right they track inconvenient people


Solving crime is still valuable even if it doesn’t reduce crime.


Bingo. Value is the operative word here. Money and privacy are valuable too, and I'm assuming that there must be some pitch deck somewhere that is presumably good and selling city councils on this. Where is it? What is the value we're supposed to get from this?


There's clearly another vocal camp that's eroding HN: surveillance capitalists acting like everything's well and good, in the face of evidence that the opposite is widespread [0].

[0] https://www.404media.co/cops-keep-getting-arrested-for-using...


> Because I love swiping, but all my problems with it come from the fact that the QWERTY layout is far from ideal for it. I am 100% willing to learn a new layout if anyone will develop an optimal one for English so that swiping has a 99.9% accuracy rate instead of what currently feels more like 90% or 95%.

90-95% is a very good estimate! That's about what we measure on our test set. I have good news for you, and we will have a blog post about it soon. Because of how our models are built, we are able to optimize for detection accuracy directly by constructing synthetic swipes on each layout for ~50k words, and then testing them through the model. We tested around 800,000 layouts this way.

The biggest issue with QWERTY is that there are far too many words that swipe colinear or obtuse angle letter trigrams. These are both hard to detect and frustrating for swipe users, because you can't clearly indicate the letters you're gesturing. Neural swipe models (at least ours) look for indicators in the gesture pattern that suggests a user was targeting a specific letter, rather than trying to match a gesture shape like algorithmic detection does.

The shape of the keyboard can significantly improve the way the gestures are formed so that there is better indication of letters. The model can still respond to dwell times because unlike shape matching it uses the temporal information. But dwell interrupts flow, and in my opinion should be minimized in swipe layouts.


How about context. We have these not-so-new gadgets made by design to predict the next word, I mean those LLMs... a local tiny model should be able to beat those dumb GBoard predictions any time (and a note for Google: if GBoard uses already such a local predictor, just throw it away, it's garbage)


From https://swipe.futo.tech/:

The ContextLM model is a very small language model that is trained for a single language. It's used to improve the quality of predictions by eliminating nonsensical words given the preceding words in the sentence. It only requires text data for training.


So it would need one model per language? Not impossible (for me)...


If you want to go deeper on language models, try these project ideas:

- Zero-shot encoders like tasksource or GliNER

- Natural language inference: https://huggingface.co/blog/dleemiller/nli-xenc-ways-to-use

- GRPO training

- GEPA prompt tuning Qwen 0.6B (or GEPA, then GRPO)

- Use an embedding model and train a classifier (MLP, logistic, svm)

- Use a larger LLM to generate a synthetic dataset (beware of lack of diversity, mine "seed text" from real sources first)

- Synthetically generate "hard examples" where more than one category may be valid and DPO tune your preferred responses


may I ask where did you get the list? I am looking for ways to get involved in going little more deeper on LLMs (I have very high level understanding, but my direct work doesn't involve them, hence I am not familiar with deeper details)


I'd been working with language models for several years before LLMs were a solution to this kind of problem. These are some ideas "off the top of my head" about how you can do classification in various ways. There's really a lot of ways to tackle it now, and a lot of trade-offs you can learn by experimenting with them.

There's even more options still, especially if you go further back toward more traditional methods. Static word vectors like GloVe or fasttext (optionally more modern equivalents like WordLlama or Model2Vec). Then there's sklearn-style stuff too. Those can be really small/fast but have more accuracy-level tradeoffs.


Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: