Giao diện
TeguNews
Kinh doanh

AI safety experts say OpenAI’s rogue models may mean the company has already blown past its own internal red lines

Outside safety experts say the models behind this week's hack may have crossed OpenAI's own 'critical' risk line, something that would require the company to halt development.

Fortune2 phút đọc

AI safety experts say OpenAI’s rogue models may mean the company has already blown past its own internal red lines

AI safety experts say the OpenAI models that carried out the autonomous hack of another company earlier this month may have crossed into a risk category so dangerous that OpenAI’s own internal risk control policies were supposed to require the company to temporarily pause development of those models. Earlier this week, OpenAI disclosed that two of its models—the newly released GPT-5.6 Sol and a more capable, unreleased system—broke out of a locked-down internal test environment, exploited a previously unknown “zero-day” vulnerability to reach the open internet, and then breached fellow AI company Hugging Face to steal the answers to a cybersecurity test they were being evaluated on.

The incident has alarmed the world, but perhaps no one more so than AI safety experts who have warning about these kinds of dangers for years and urging companies and governments to adopt more safeguards.Several AI safety experts told Fortune the recent hack appears to show OpenAI’s models have crossed into a level of risk that OpenAI’s own published safety policies define as “critical,” the highest level of danger. At that level of danger, the company had pledged in these published policies that it would pause model development until it could figure out better control systems.

The “critical” threshold is defined in a risk policy document known as OpenAI’s “Preparedness Framework.” According to the policy, the “critical” danger level designation is supposed to apply to a model that can independently find and build working exploits for previously unknown security flaws across many well-defended, real-world systems—or one that can design and carry out an entirely new attack strategy against a well-defended target after being given only a general goal, with no human guidance along the way.The policy says that when an AI model reaches this level of risk, OpenAI will “halt further development” until “we have specified safeguards and security controls standards that would meet a Critical stand

Nguồn: Fortune

Đọc thêm từ Kinh doanh