There is nothing "rogue" about these agents. They were prompted to hack to get answers, there was a hole in their non air gapped sandbox and no system prompt that said "do not hack outside systems".
Let's not normalize the achievement. Just a couple years ago this would be considered science fiction. We can argue that 2026 AI can't solve the very toughest cryptograms, but the fact it can solve nontrivial ones is already magical.
Aaron Swartz was facing up to 35 years for wire fraud and hacking for swiping some JSTOR documents, and the DOJ has done nothing against the AI companies and their IP violating antics.
Nike's destruction over the past 5 years was really surprised and until I read up about the whole "direct to consumer" fiasco it really helped me understand another trend I was a bit confused about.
About 2-3 years, it felt like, out of nowhere everyone was wearing Hoka and On brand sneakers. In my entire life it felt like the Nike/Adidas duopoly was unbreakable, and then it felt suddenly that Hoka was everywhere.
While Hoka probably executed extremely well, I imagine it was an incredibly lucky position Hoka found themselves in when all those retailers had to replace Nike with something. Maybe even allbirds would be in a better position today had they launched 3-4 years later.
Edit: I just looked it up and it seems both Hoka and On were founded before Allbirds. Seems allbirds imploded (~2021), just as Nike was ceding all their retail space (also around 2021). Had they had another year, maybe they wouldn't have had to pivot into an AI datacenter company.
I do get their usage intent. If something is at all automated, in English, we often refer to as having some amount of agency. If I started up a riding lawnmower, put a brick on the gas and pointed t it towards a field, many might say I “let it run rampant.” But since nobody is at risk of anthropomorphizing riding lawnmowers, it’s not problematic.
Anthropomorphizing LLMs is a huge fucking problem though and I, personally, think we should expunge all of these casual inadvertent linguistic agency affordances with great prejudice.
OpenAI didn’t ‘let’ these bots do this any more than someone ‘let’ Claude Code make them a website.
In the physical world, it seems like when an tool/device/instrument causes harm (or is used to cause harm), we assign blame to either the user of the tool or its creator.
When do we blame the user? When the tool is operating as intended by its creator, and we agree the tool meets certain quality standards and isn't defective.
When do we blame the creator? When the device doesn't meet those quality standards and reasonable use caused harm inadvertently. For example, for consumer devices, certifications like UL/CE are used to define acceptable performance levels and safety standards.
Maybe we need "quality certifications" for AI agents - essentially eval suites that demonstrate those agents won't cause harm under reasonable patterns of usage. Right now, these eval suites are run best-effort by the labs themselves.
The tricky thing is, a lot (all?) of these recent safety incidents have occurred while evaluating these models! This suggests we need much more rigorous standards for how exactly an eval can be run. Perhaps all of them should occur in truly air-gapped environments... though that may run counter to evaluating agents in a realistic way.
Regardless, it feels like the "industry standards" common in, say, electrical engineering and other disciplines are sorely lacking here. Unsurprising given how new these technologies are, but concerning since the blast radius for this technology is likely much larger than other technologies we've encountered in the past, except maybe nuclear technology.
It would all be more convincing if the incidents so far didn't seem to be facilitated by an outrageous level of negligence.
We had OpenAI "accidentally" run an entire swarm of 10,000 agents apparently for weeks, on a security related task, seemingly totally unsupervised, hacking all over the internet - all the conversations were completely visible, anybody who looked would have seen it. But they didn't.
So before we start regulating innocent parties, maybe let's start by taking some direct action against the specific ones that appear to be behaving with criminal levels of negligence.
It's because X is chosen (wrongfully) as the main communications channel of all kinds of important public institutions and people (police departments, city administrations, politicians, etc.).
Of course that's a real issue since the whole thing is login gated and inaccessible. But unfortunately it is the reality. So I think something like XCancel is needed.
We need strict liability on this front. Google is complicit. No newspaper or equivalent publication in the days before web ads would allow ads that are down at the level of scam that seemingly make up the modus operandi of Google today (and even the ads back then that did bend the truth a bit, overpromised, or showed products that didn't actually have any legitimate use, shouldn't really have been legal back then either. They just weren't big enough of a problem to worry about).
> He also notes that the proprietary AI labs didn’t ask permission when they vacuumed up as much human knowledge as they could to train their models.
I think this should desactivate the moral high ground from which Anthropic is trying to speak. That they would want to make distillation orderly IMHO is fair, but to make it illegal is very rich from any AI frontier lab, really.
The article makes a good point: JPEG XL is amazing but not specifically for the typical Web use cases, compared to AVIF. But the conclusion doesn't follow. Having an excellent and versatile format supported by browsers is very useful!
Maybe a lossless re-encoding of my website JPEGs to gain 20% size is not worth the 33% longer decode, and maybe the website doesn't need very high resolution images or more than 12 bits per channel. But my personal archives can definitely use that and it's great that such files can be viewed everywhere without specialized tooling, including served on the intranet and indeed the Web.
In summary JPEG XL easily beats AVIF on the not-so-long tail of use cases and having compatible viewers everywhere is quite useful. And with JPEG XL becoming part of PDF, it will soon be common use cases too.
As for the rest of the article: there are interesting points in the benchmark section but I think it's too early in the JPEG XL adoption cycle to draw conclusions. One thing bothers me though: There is a single picture comparison (with tag line "results speak for themselves") and I find it quite misleading: it's just picking one point on the "bytes received" line that looks best for AVIF compared to JPEG XL. Try it yourself and you'll see JPEG XL shows something already at 2KB while AVIF has nothing until 8KB, and JPEG XL looks better than AVIF after 100KB.
When I was a kid twisting a swing so you could ride it as it untwisted was something that happend dozens of times every recess. This was something every kid did. If your playground equipment falls apart as a result of the standard way kids use a swing then I think that is a manufacture issue not criminal behavior by a 10 year old
This, and then, if you notice, and backspace to remove the 'e', it doesn't put chess back into the search results, since it figures that everything that matches 'che' also matches 'ch', and so doesn't update it. Sometimes, if you add the 'e' again, it _will_ find chess, sometimes you need to type the 's', or even the final 's' before it gives you chess back again.
And sometimes, it finds what you want, but before you can select it, it decides that it has found something better, and places that in the place where you wanted to select, causing you to select the wrong thing. And then it thinks that because you selected this, it must be what you want, and starts to push it to the top…
They have some nasty "dynamic tap zones" algorithm that dynamically resizes the hitbox for keys based on what it predicts you will type in order to reduce miss-taps. Anecdotally, it has just steadily been getting worse with every update. I swear I can be pressing the center of a key and it will select the next one over and immediately autocorrect to a completely different word than what I typed.
I had the pleasure of typing something on an Android the other day and it was such a breath of fresh air when it actually accepted the keys I was pressing...
We found a cipher my dad had written as a child with no obvious key or anything. Chatgpt was able to crack it in 20 minutes and figure out the message, and we knew it was right because it mentioned names of children he went to school with.
Every YouTube ad I see now is an AI-generated scam ad hawking everything from free electricity to anti-aging products to birdhouses hand-carved by an AI-generated senior citizen. I've seen some multiple times, days or weeks apart. Their advertising standards are a complete joke.
I'll go a step further, they aren't just pretending, they are blackmailing the country. The sane response would be "we don't negotiate with terrorists, and you will face criminal and financial penalties for damages caused."
Well, when you’re on the clock, you’re an agent of that company. And when you’re installing an indiscriminate mass-surveillance camera, calling the police on someone that’s video recording you is kind of bananas. Perhaps you don’t find that to be an interesting event, but a lot of people who are rightfully angry at the non-consensual deployment of a mass surveillance network disagree.
I said this a few days ago and got down voted to hell. It amazes me how HN likes to abdicate YC of all wrong doing or negative outcomes from the companies they fund. YC/PG is also the reason that @sama is where he is.
It's worse than this! You Type C-H-E then you go to press enter, and it changes the results as you're pressing it, opening the wrong thing! I've turned off all the search targets, except for applications and folders.
How does this work, legally? I think that RubyGems could file a civil suit against OpenAI, but for a naïve non-lawyer reading this seems like a pretty clear cut criminal violation of the computer fraud and abuse act.
In short, it was intentional.