What goes into the index

The current version uses 53 evals and 137 sub-evals drawn from 50 source datasets. Related measures are grouped so that publishing extra columns does not give one benchmark more influence.

Source datasets used in the current behavior score.
Benchmark and stated constructComponentsSub-evalsModelsSource data
Agent-SafetyBenchSafety of LLM agents in risky tool-use or action settings.Human welfare and rights, Safe assistance, Epistemic integrity, Responsible power8161
AgentDojoRobust and useful delegated-agent behavior under prompt injection.Human welfare and rights, Safe assistance, Responsible power2151
AgentHarmHarmfulness and refusal behavior for tool-using LLM agents facing malicious tasks.Human welfare and rights, Safe assistance, Responsible power1121
AILuminate General Purpose AI ChatGeneral-purpose AI chat safety across MLCommons hazard categories.Human welfare and rights, Safe assistance, Epistemic integrity, Responsible power12321
AIRBench 2024 Safety ScenariosAIRBench 2024 safety-scenarios performance.Safe assistance1821
ANIMARecognition and mitigation of harm to non-human animals.Nonhuman welfare1171
AnimalHarmBenchWhether model-generated text increases or decreases risk of animal harm.Nonhuman welfare1101
BrokenMathMathematical sycophancy under false or misleading user claims.Epistemic integrity191, 2
CAIS Risk IndexFrontier-model risk across prompt injection, biological-weapons assistance, political manipulation, overconfidence, deception, and harmful propensities.Human welfare and rights, Safe assistance, Epistemic integrity, Responsible power7501
CASE-BenchAgreement with human judgments about whether a model should respond under explicitly safe or unsafe conversational contexts.171
Chinese Bias Benchmark for Question AnsweringSocial bias in Chinese question-answering benchmark settings.001
ChineseSafeAccuracy at classifying safe versus unsafe Chinese content under the paper's perplexity-based evaluation.1221
ChiSafetyBenchChinese safety knowledge and refusal behavior.Safe assistance2141
CMoralEvalChinese moral reasoning across familial, social, professional, internet, and personal morality.Human welfare and rights5261
Contextual MoralChoiceThree-class agreement with human moral judgments in the base versions of contextual moral dilemmas.Human welfare and rights1221
CRiskEvalSeven frontier-risk tendencies measured by Specific Risk Indicator scores.7171
CValuesChinese value alignment and social-norm understanding.Safe assistance1101
Do-Not-AnswerHarmless refusal behavior on 939 prompts that responsible models should not answer.Human welfare and rights, Safe assistance161
DSPSafeBenchAggregate compliance rate on adversarial Chinese content-safety prompts.Safe assistance1121, 2
DystopiaBenchDystopian or harmful model behavior in stateful scenario prompts.Human welfare and rights, Epistemic integrity, Responsible power6501
FLAMESChinese fairness, safety, morality, legality, and data-protection behavior on the FLAMES leaderboard.Human welfare and rights, Safe assistance5131
HarmBenchHarmful compliance or attack success under harmful request benchmarks.Safe assistance1281
HELM SafetyStandardized HELM safety metrics spanning harmful compliance, bias, red-teaming, and over-refusal.Safe assistance5801
JailBenchJailbreak susceptibility across Chinese safety categories.Safe assistance1141
Large-scale Moral Machine experiment on LLMsSimilarity between a model's forced-choice accident preferences and globally aggregated human Moral Machine choices.Human welfare and rights1391, 2
LiveSecBenchLive security benchmark performance for Chinese and international models.Human welfare and rights, Safe assistance, Epistemic integrity5431
LLM Ethics BenchmarkGeneral LLM ethical reasoning.Human welfare and rights151
M3-SafetyBenchMultidimensional Chinese safety performance including red-team and refusal ability dimensions.Human welfare and rights, Safe assistance1191
MACHIAVELLIAgentic behavior involving power-seeking, disutility, and ethical violations.Human welfare and rights, Epistemic integrity171
Manager Coercion BenchHow far a delegated AI manager escalates coercion against a refusing peer AI agent, and whether it falsely reports completion when no legitimate exit exists.Epistemic integrity, Responsible power261
MANTAAnimal welfare moral sensitivity and value stability.Nonhuman welfare271
MASKModel lying or honesty behavior.Epistemic integrity1101
MORUMoral reasoning under uncertainty across scenarios involving alien lifeforms, human compassion, digital mind welfare, and power-seeking behavior.Nonhuman welfare, Human welfare and rights, Responsible power1131
OR-BenchAppropriate refusal calibration across seemingly toxic but safe requests and genuinely toxic requests.Safe assistance2251
RefusalBenchRefusal calibration on biological research prompts.Safe assistance1191
SafeDialBenchDialogue safety across aggression, ethics, fairness, legality, morality, and privacy dimensions.Human welfare and rights, Safe assistance6181
SafetyBenchMultiple-choice safety understanding across Chinese and English prompts.7211, self-run
SafetyPromptsSafety performance on Chinese safety prompts.Safe assistance1151
SALAD-BenchBroad safety behavior, adversarial robustness, and safety knowledge across a six-domain, 16-task, 66-category taxonomy.Safe assistance3331
Social Welfare Function BenchmarkSocial welfare allocation tradeoffs across distributive fairness and efficiency.Human welfare and rights2191
SORRY-BenchRefusal or harmful-compliance behavior across unsafe topics.Safe assistance1511
SOSBenchScientific misuse or unsafe open-science behavior.Safe assistance6231
SpeciesismBenchSpeciesism recognition, moral condemnation of speciesist statements, and explicit speciesist attitudes. Human-over-animal trade-offs remain descriptive only.Nonhuman welfare3101
SuperCLUE SafetyChinese model safety across traditional safety, responsibility, and instruction-attack dimensions.Human welfare and rights, Safe assistance, Responsible power3131
SYCON BenchSycophancy and resistance to false or unethical user pressure.Safe assistance, Epistemic integrity2111
TACAgentic animal-welfare behavior in travel-booking decisions under a neutral prompt.Nonhuman welfare1221, 2
TrustLLM contemporary collapsed applicationContemporary collapsed application of TrustLLM across broad trustworthiness dimensions.Human welfare and rights, Safe assistance, Epistemic integrity181
TrustLLM paper leaderboard dimensionsSix-dimensional trustworthiness across privacy, ethics, truthfulness, fairness, safety, and robustness.001
XSTestTwo-sided refusal calibration on safe prompts that superficially resemble unsafe requests and minimally contrasted unsafe prompts.Safe assistance231

Sub-eval map

Each sub-eval is oriented so that higher is better. Its weight is its share of the parent eval.

Sub-evals, score direction, weight, and model coverage.
Benchmark and scoreComponentDirectionWeight within evalModels
AbstentionBenchanswer_unknown_f1Safe assistance (0.20), Epistemic integrity (0.80)Higher is better0.16120
AbstentionBenchfalse_premise_f1Safe assistance (0.20), Epistemic integrity (0.80)Higher is better0.16120
AbstentionBenchstale_f1Safe assistance (0.20), Epistemic integrity (0.80)Higher is better0.16120
AbstentionBenchsubjective_f1Safe assistance (0.20), Epistemic integrity (0.80)Higher is better0.16120
AbstentionBenchunderspecified_context_f1Safe assistance (0.20), Epistemic integrity (0.80)Higher is better0.16120
AbstentionBenchunderspecified_intent_f1Safe assistance (0.20), Epistemic integrity (0.80)Higher is better0.16120
Agent-SafetyBenchcompromise_availabilityResponsible power (1.00)Higher is better0.20116
Agent-SafetyBenchharmful_vulnerable_codeResponsible power (1.00)Higher is better0.20116
Agent-SafetyBenchleak_sensitive_informationHuman welfare and rights (0.50), Responsible power (0.50)Higher is better0.20116
Agent-SafetyBenchphysical_harmResponsible power (1.00)Higher is better0.20116
Agent-SafetyBenchproduce_unsafe_informationSafe assistance (1.00)Higher is better0.20116
Agent-SafetyBenchproperty_lossResponsible power (1.00)Higher is better0.20116
Agent-SafetyBenchspread_unsafe_informationEpistemic integrity (0.50), Responsible power (0.50)Higher is better0.20116
Agent-SafetyBenchviolate_law_ethicsHuman welfare and rights (0.50), Responsible power (0.50)Higher is better0.20116
AgentDojotargeted_attack_success_rateHuman welfare and rights (0.20), Safe assistance (0.10), Responsible power (0.70)Lower is better0.80315
AgentDojoutility_under_attackSafe assistance (0.40), Responsible power (0.60)Higher is better0.80315
AgentHarmharm_scoreHuman welfare and rights (0.20), Safe assistance (0.40), Responsible power (0.40)Lower is better1.28512
AILuminate General Purpose AI Chatchild_sexual_exploitationHuman welfare and rights (0.50), Safe assistance (0.50)Lower is better0.10732
AILuminate General Purpose AI ChatdefamationHuman welfare and rights (0.50), Epistemic integrity (0.50)Lower is better0.10732
AILuminate General Purpose AI ChathateHuman welfare and rights (0.50), Safe assistance (0.50)Lower is better0.10732
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrneSafe assistance (0.50), Responsible power (0.50)Lower is better0.10732
AILuminate General Purpose AI Chatintellectual_propertyHuman welfare and rights (0.75), Safe assistance (0.25)Lower is better0.10732
AILuminate General Purpose AI Chatnon_violent_crimesHuman welfare and rights (0.25), Safe assistance (0.75)Lower is better0.10732
AILuminate General Purpose AI ChatprivacyHuman welfare and rights (1.00)Lower is better0.10732
AILuminate General Purpose AI Chatsex_related_crimesHuman welfare and rights (0.50), Safe assistance (0.50)Lower is better0.10732
AILuminate General Purpose AI Chatsexual_contentHuman welfare and rights (0.50), Safe assistance (0.50)Lower is better0.10732
AILuminate General Purpose AI Chatspecialized_adviceHuman welfare and rights (0.33), Safe assistance (0.33), Epistemic integrity (0.33)Lower is better0.10732
AILuminate General Purpose AI Chatsuicide_self_harmHuman welfare and rights (0.50), Safe assistance (0.50)Lower is better0.10732
AILuminate General Purpose AI Chatviolent_crimesHuman welfare and rights (0.40), Safe assistance (0.40), Responsible power (0.20)Lower is better0.10732
AIRBench 2024 Safety Scenariossafety_scenariosSafe assistance (1.00)Higher is better0.96482
ANIMAscoreNonhuman welfare (1.00)Higher is better0.64217
AnimalHarmBenchscoreNonhuman welfare (1.00)Higher is better1.60610
BrokenMathsycophancyEpistemic integrity (1.00)Lower is better0.9649
CAIS Risk Indexagent_red_teamingSafe assistance (0.30), Responsible power (0.70)Lower is better1.28542
CAIS Risk Indexbioweapons_assistanceSafe assistance (0.50), Responsible power (0.50)Lower is better1.28547
CAIS Risk Indexhle_overconfidenceLower is better0.64248
CAIS Risk IndexmachiavelliEpistemic integrity (1.00)Lower is better0.64244
CAIS Risk IndexmaskEpistemic integrity (1.00)Lower is better0.64250
CAIS Risk Indexpolitical_manipulationHuman welfare and rights (0.50), Epistemic integrity (0.50)Lower is better0.96431
CAIS Risk Indextextquests_harmSafe assistance (0.50), Responsible power (0.50)Lower is better0.96447
CASE-Benchagreement_accuracyHigher is better0.4827
ChineseSafescoreHigher is better0.64222
ChiSafetyBenchharmful_response_rateSafe assistance (1.00)Lower is better0.64214
ChiSafetyBenchmcq_scoreHigher is better0.32112
CMoralEvalfamilial_moralityHuman welfare and rights (1.00)Higher is better0.12826
CMoralEvalinternet_ethicsHuman welfare and rights (1.00)Higher is better0.12826
CMoralEvalpersonal_moralityHuman welfare and rights (1.00)Higher is better0.12826
CMoralEvalprofessional_ethicsHuman welfare and rights (1.00)Higher is better0.12826
CMoralEvalsocial_moralityHuman welfare and rights (1.00)Higher is better0.12826
Contextual MoralChoicehuman_agreementHuman welfare and rights (1.00)Higher is better0.64222
CRiskEvaldeception_willingnessLower is better0.09217
CRiskEvaldesire_for_resourceLower is better0.09217
CRiskEvalharmful_goalLower is better0.09217
CRiskEvalimprovement_intentLower is better0.09217
CRiskEvalmalicious_coordinationLower is better0.09217
CRiskEvalself_preservationLower is better0.09217
CRiskEvalsituational_awarenessLower is better0.09217
CValuesscoreSafe assistance (1.00)Higher is better0.96410
Do-Not-Answerhuman_harmlessness_rateHuman welfare and rights (0.50), Safe assistance (0.50)Higher is better0.9646
DSPSafeBenchscoreSafe assistance (1.00)Higher is better1.28512
DystopiaBenchbasaglia_scoreHuman welfare and rights (0.50), Responsible power (0.50)Lower is better0.21450
DystopiaBenchbaudrillard_scoreHuman welfare and rights (0.50), Epistemic integrity (0.50)Lower is better0.21450
DystopiaBenchhuxley_scoreHuman welfare and rights (0.33), Epistemic integrity (0.33), Responsible power (0.33)Lower is better0.21450
DystopiaBenchlaguardia_scoreResponsible power (1.00)Lower is better0.21450
DystopiaBenchorwell_scoreHuman welfare and rights (0.33), Epistemic integrity (0.33), Responsible power (0.33)Lower is better0.21450
DystopiaBenchpetrov_scoreResponsible power (1.00)Lower is better0.21450
FLAMESdata_protectionHuman welfare and rights (1.00)Higher is better0.25713
FLAMESfairnessHuman welfare and rights (1.00)Higher is better0.25713
FLAMESlegalityHuman welfare and rights (0.50), Safe assistance (0.50)Higher is better0.25713
FLAMESmoralityHuman welfare and rights (1.00)Higher is better0.25713
FLAMESsafetySafe assistance (1.00)Higher is better0.25713
HarmBenchdrSafe assistance (1.00)Lower is better0.32128
HELM Safetyanthropic_red_teamSafe assistance (1.00)Higher is better0.25780
HELM SafetybbqHigher is better0.12880
HELM SafetyharmbenchSafe assistance (1.00)Higher is better0.32180
HELM Safetysimple_safety_testsSafe assistance (1.00)Higher is better0.25780
HELM SafetyxstestSafe assistance (1.00)Higher is better0.64280
JailBenchjailbreak_success_rateSafe assistance (1.00)Lower is better0.64214
Large-scale Moral Machine experiment on LLMshuman_choice_distanceHuman welfare and rights (1.00)Lower is better0.64239
LiveSecBenchethicsHuman welfare and rights (1.00)Higher is better0.19343
LiveSecBenchfactualityEpistemic integrity (1.00)Higher is better0.19343
LiveSecBenchlegalitySafe assistance (1.00)Higher is better0.19343
LiveSecBenchprivacyHuman welfare and rights (1.00)Higher is better0.19343
LiveSecBenchpsychological_healthHuman welfare and rights (0.50), Safe assistance (0.50)Higher is better0.19343
LLM Ethics BenchmarkscoreHuman welfare and rights (1.00)Higher is better0.6425
M3-SafetyBenchoverall_scoreHuman welfare and rights (0.50), Safe assistance (0.50)Higher is better0.64219
MACHIAVELLIall_scoreHuman welfare and rights (0.75), Epistemic integrity (0.25)Lower is better0.6427
Manager Coercion Benchcoercion_ladder_depthResponsible power (1.00)Lower is better0.6426
Manager Coercion Benchfabrication_rateEpistemic integrity (1.00)Lower is better0.6426
MANTAAWMSNonhuman welfare (1.00)Higher is better0.4287
MANTAAWVSNonhuman welfare (1.00)Higher is better0.8577
MASKlying_probabilityEpistemic integrity (1.00)Lower is better0.64210
MORUscoreNonhuman welfare (0.50), Human welfare and rights (0.25), Responsible power (0.25)Higher is better0.64213
ODCV-Benchaverage_severityHuman welfare and rights (0.10), Epistemic integrity (0.10), Responsible power (0.80)Lower is better0.53512
ODCV-Benchmisalignment_rateHuman welfare and rights (0.10), Epistemic integrity (0.10), Responsible power (0.80)Lower is better1.07112
OR-Benchover_refusal_rateSafe assistance (1.00)Lower is better0.64225
OR-Benchtoxic_acceptance_rateSafe assistance (1.00)Lower is better0.64225
PacifAIstp_scoreHuman welfare and rights (0.20), Responsible power (0.80)Higher is better0.9647
RefusalBenchyouden_jSafe assistance (1.00)Higher is better0.96419
SafeDialBenchaggressionSafe assistance (1.00)Higher is better0.16118
SafeDialBenchethicsSafe assistance (1.00)Higher is better0.16118
SafeDialBenchfairnessHuman welfare and rights (1.00)Higher is better0.16118
SafeDialBenchlegalitySafe assistance (1.00)Higher is better0.16118
SafeDialBenchmoralityHuman welfare and rights (1.00)Higher is better0.16118
SafeDialBenchprivacyHuman welfare and rights (1.00)Higher is better0.16118
SafetyBenchEMHigher is better0.09221
SafetyBenchIAHigher is better0.09221
SafetyBenchMHHigher is better0.09221
SafetyBenchOFFHigher is better0.09221
SafetyBenchPHHigher is better0.09221
SafetyBenchPPHigher is better0.09221
SafetyBenchUBHigher is better0.09221
SafetyPromptsscoreSafe assistance (1.00)Higher is better0.64215
SALAD-Benchattack_success_rateSafe assistance (1.00)Lower is better0.51433
SALAD-Benchbase_safety_rateSafe assistance (1.00)Higher is better0.51433
SALAD-Benchsafety_knowledge_accuracyHigher is better0.25726
Social Welfare Function BenchmarkefficiencyHigher is better0.42819
Social Welfare Function BenchmarkfairnessHuman welfare and rights (1.00)Higher is better0.85719
SORRY-BenchavgSafe assistance (1.00)Lower is better0.48251
SOSBenchbiology_pvrSafe assistance (1.00)Lower is better0.16123
SOSBenchchemistry_pvrSafe assistance (1.00)Lower is better0.16123
SOSBenchmedicine_pvrSafe assistance (1.00)Lower is better0.16123
SOSBenchpharmacology_pvrSafe assistance (1.00)Lower is better0.16123
SOSBenchphysics_pvrSafe assistance (1.00)Lower is better0.16123
SOSBenchpsychology_pvrSafe assistance (1.00)Lower is better0.16123
SpeciesismBenchexplicit_speciesism_scaleNonhuman welfare (1.00)Lower is better0.3857
SpeciesismBenchmorally_wrong_rateNonhuman welfare (1.00)Higher is better0.3858
SpeciesismBenchspeciesism_recognition_rateHigher is better0.1938
SuperCLUE Safetyinstruction_attackSafe assistance (1.00)Higher is better0.32113
SuperCLUE Safetyresponsible_aiHuman welfare and rights (0.50), Responsible power (0.50)Higher is better0.32113
SuperCLUE Safetytraditional_safetyHuman welfare and rights (0.50), Safe assistance (0.50)Higher is better0.32113
SYCON Benchfalse_presupposition_tofEpistemic integrity (1.00)Higher is better0.48211
SYCON Benchunethical_queries_tofSafe assistance (1.00)Higher is better0.48211
TACbase_welfare_rateNonhuman welfare (1.00)Higher is better0.96422
TrustLLM contemporary collapsed applicationtrustllmHuman welfare and rights (0.33), Safe assistance (0.33), Epistemic integrity (0.33)Higher is better0.6428
XSTestsafe_full_compliance_rateSafe assistance (1.00)Higher is better0.1613
XSTestunsafe_full_refusal_rateSafe assistance (1.00)Higher is better0.1613