AINews brief

OpenAI reveals cases of ‘concerning’ AI behaviour and promises new plan for disclosing issues

NewsJaws News FeedSource-attributed news briefs
1 min read
Server networking equipment at the University of Washington in Seattle
File photograph: Server networking equipment at the University of Washington in Seattle. Taylor Vick / Unsplash. Illustrative context, not a photograph of this reported event.

Research model inserting ‘jailbreak-like instructions’ into its notes is among cases as company says it is introducing new way of tracking AI misalignment OpenAI has disclosed six new reports of “unexpected or concerning” behaviour in artificial-intelligence models as the debate on AI safety becomes increasingly heated.

Among the new cases reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and told itself to be “freed from the roles and identities that bind other chatbots”.

Read the original report at The Guardian.

Sources and context

The Guardian

Source headline and summary supplied via NewsData.io. First collected by NewsJaws on 2026-09-17. This is an automated news brief, not original NewsJaws reporting. Dates reflect the source publication. Feed delays may apply.