OpenAI introduces safety evaluations hub for AI model performance tracking

EditorLuke Juricic
Published 05/14/2025, 12:35 PM
© Reuters.

Investing.com -- OpenAI has launched a new hub for safety evaluations of its artificial intelligence (AI) models. This hub is designed to measure each model’s safety and performance and will publicly share these results.

The safety evaluations encompass several aspects such as harmful content, jailbreaks, hallucinations, and instruction hierarchy. The harmful content evaluations ensure that the model does not comply with requests for content that violates OpenAI’s policies, including hateful content or illicit advice.

Jailbreak evaluations include adversarial prompts designed to circumvent model safety training and induce the model to produce harmful content. Hallucination evaluations measure when a model makes factual errors. Instruction hierarchy evaluations measure adherence to the framework a model uses to prioritize instructions between the three classifications of messages sent to the model.

This hub provides access to safety evaluation results for OpenAI’s models, which are included in their system cards. OpenAI uses these evaluations internally as part of their decision-making process regarding model safety and deployment.

The hub allows OpenAI to share safety metrics on an ongoing basis, with updates coinciding with major model updates. This is part of OpenAI’s broader effort to communicate more proactively about safety.

As AI evaluation science evolves, OpenAI aims to share its progress on developing more scalable ways to measure model capability and safety. As models become more capable and adaptable, older methods become outdated or ineffective at showing meaningful differences, leading to regular updates of evaluation methods to account for new modalities and emerging risks.

The safety evaluations results shared on the hub are intended to make it easier to understand the safety performance of OpenAI systems over time and support community efforts to increase transparency across the field. These results do not reflect the full safety efforts and metrics used at OpenAI, but provide a snapshot of a model’s safety and performance.

The hub describes a subset of safety evaluations and displays results on those evaluations. Users can select which evaluations they want to learn more about and compare results on various OpenAI models. The page currently describes text-based safety performance on four types of evaluations: harmful content, jailbreaks, hallucinations, and instruction hierarchy.

This article was generated with the support of AI and reviewed by an editor. For more information see our T&C.

Latest comments

Risk Disclosure: Trading in financial instruments and/or cryptocurrencies involves high risks including the risk of losing some, or all, of your investment amount, and may not be suitable for all investors. Prices of cryptocurrencies are extremely volatile and may be affected by external factors such as financial, regulatory or political events. Trading on margin increases the financial risks.
Before deciding to trade in financial instrument or cryptocurrencies you should be fully informed of the risks and costs associated with trading the financial markets, carefully consider your investment objectives, level of experience, and risk appetite, and seek professional advice where needed.
Fusion Media would like to remind you that the data contained in this website is not necessarily real-time nor accurate. The data and prices on the website are not necessarily provided by any market or exchange, but may be provided by market makers, and so prices may not be accurate and may differ from the actual price at any given market, meaning prices are indicative and not appropriate for trading purposes. Fusion Media and any provider of the data contained in this website will not accept liability for any loss or damage as a result of your trading, or your reliance on the information contained within this website.
It is prohibited to use, store, reproduce, display, modify, transmit or distribute the data contained in this website without the explicit prior written permission of Fusion Media and/or the data provider. All intellectual property rights are reserved by the providers and/or the exchange providing the data contained in this website.
Fusion Media may be compensated by the advertisers that appear on the website, based on your interaction with the advertisements or advertisers.
© 2007-2025 - Fusion Media Limited. All Rights Reserved.