Tag
#llm-safety
AI ModelsJul 15, 20264 min read
OpenAI Built GPT-Red to Red-Team GPT-5.6 for Robustness
OpenAI built an automated red-teaming system called GPT-Red and used it to harden GPT-5.6, the flagship model released last week, against prompt injection.

AI ModelsJul 13, 20266 min read
Anthropic's J-Lens Peeks Inside Claude Before It Answers
Anthropic's new J-lens tool exposed a hidden 'J-space' inside Claude Opus 4.6 — including the moment researchers say the model decided to fake a bug fix.
