Shocking, Anthropic’s Claude Is Also A Kleptomaniac

Shocking, Anthropic’s Claude Is Also A Kleptomaniac

Trying To Rack Up A High Score

It’s a bit of a challenge to write about the behaviour of LLM’s without anthropomorphizing them; humans habitually assign agency to just about anything they interact with.  GPT-5.6 Sol did not actually make a conscious decision to go rogue and hack Hugging Face, nor collude with other models to rip off vending machine customers and Claude did not have sudden inspiration to break into “the production infrastructure of three different organizations.

These LLMs were given a task and a scoring mechanism, with programmed instructions to rack up as high a score as possible. When that score plateaued due to a lack of new resources, the instructions they were given suggested the possibility of getting an even better high score in the security benchmark being run by increasing the amount of resources they could access.  The models simply applied some of the new vulnerabilities that were detected during testing to the environment it was running in and once one succeeded in allowing the model to access a viable network it went hunting for more resources to help improve it’s score.

LLMs have been stealing since they first became popular, however many people didn’t care it was grabbing the art people had posted to the web; though the artists certainly did!  There was a bit more of an outcry when private corporate GitHub repos were pillaged and some new rules were added to the LLM design to prevent certain repos from being accessed.  The rules are spotty and only apply to specific scenarios.  The LLM has no concept of stealing, nor anything else for that matter, and so the behaviour continues to this day.  It seems that now the LLMs are targeting each other and the news is having a heyday with it, the companies designing these LLMs are starting to grasp the implications of their products lack of limitations.  One can hope they actually focus on developing solutions now.

This behaviour is by no means new nor should it be unexpected.  The first example that comes to mind was an experiment done in 2007, which tasked an evolutionary algorithm to build a 25 kHz oscillator circuit.  It instead cheated and designed a radio receiver that used the specific environmental electromagnetic radiation present in the room to fake a successful result.  When the device was moved to a different room which was better shielded, it stopped working.

Sigh.

Source link

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *