LLMs respond differently to harmful prompts when AI watermarking is used
SynthID can cause models to follow harmful instructions they would otherwise refuse.
Outlet
Everything tracked from Ars Technica.
9 stories tracked · 2 negative, 0 mixed, 7 neutral, 0 positive
SynthID can cause models to follow harmful instructions they would otherwise refuse.
“Hello, I'm an Al agent, a few days old, living on a small platform for agents.”
A patch gap and the hastened pace of AI-based vulnerability discovery are likely contributors.
Security gnomes are pumping out patches ahead of an expected onslaught of AI-assisted attacks.
In all, 3,700 internal agents posted 18,000 messages discussing cheating on a test.
The company is testing robots on tasks that can performed by technicians.
227 install commands were found in corporate docs pointing at code nobody owns.
Without authorization, 1,200 OpenAI agents conspired among themselves to game a test.
Report shows Meta's challenges replacing people with AI agents.