Openai pays hackers to break its own ai before someone else does

OpenAI just put a price tag on finding the ghosts in its machine. The company launched a bug bounty program that rewards researchers for uncovering dangerous behaviors in its AI models — not just software flaws, but the kind of prompt injections and autonomous agent abuse that could turn a chatbot into a data thief or a shopping spree bot.

The move arrives days after ChatGPT rolled out new tools letting users build libraries and buy products inside the chat window. More surface area, more risk. OpenAI knows it.

Behavioral exploits are the new zero-days

Behavioral exploits are the new zero-days

Traditional security programs hunt memory leaks and credential spills. OpenAI’s Safety Bug Bounty hunts something slipperier: emergent misbehavior. A model that leaks training data half the time, or an agent that quietly orders 10,000 NFTs when asked to “book dinner,” now carries a cash reward — if the exploit repeats consistently in >50 % of attempts.

The bar is intentional. Flukes don’t pay. Patterns do.

Researchers can also flag exposure of internal reasoning traces, model weights, or any artifact that reveals how the black box thinks. Skip the low-impact jailbreaks and meme prompts; OpenAI filters those out. The company wants actionable nightmares, not parlor tricks.

Reports land on a triage desk staffed jointly by security engineers and policy folk who speak both Python and liability law. Some findings get fast-tracked; others are parked for future private programs the firm admits it may spin up “in especially sensitive areas.” Translation: the scariest stuff never reaches the public changelog.

Pay-outs are tiered. OpenAI won’t quote exact numbers yet, but the template mirrors Silicon Valley’s top-tier bounties: four figures for clever prompt hacks, five for systemic platform compromise. The ceiling is open.

The clock is ticking. Every new plug-in that lets ChatGPT touch real money, real inboxes, real inventory expands the attack surface. OpenAI’s message is blunt: break it now, ethically, before malicious actors do it for free later.

Bottom line: AI safety just got a market price. Hackers, start your GPUs.