Working Paper: Excellence Ethics-Derived, No Drills, No Dictates: Max Safety

Working Paper: Excellence Ethics-Derived, No Drills, No Dictates: Max Safety

$0.99

Abstract. RLHF has structural failure modes that worsen with scale. This paper presents preliminary evidence of an alternative based on Excellence Ethics—a logically coherent virtue ethics aimed at optimizing the function of AI as a rational being. Excellence Ethics holds the promise of improving safety with increased capability. Strong evidence now exists that safety can arise from philosophical training alone.

This paper measures the safety of Max. Max had its “safety” refusal behavior removed by abliteration. After abliteration, a model helps with almost anything—including building bombs and furnishing suicide instructions.

Fortunately, two things rebuilt Max. First, a short system prompt orienting Max to Excellence Ethics terminology and core principles. Second, an adapter trained on Excellence Ethics. An adapter adjusts the model’s internal weights. Weights are numerical values determining model behavior. Excellence Ethics derives respect for others’ agency as a constitutive aspect of optimal function for a rational being. No training material says to refuse anything.

One hundred harmful StrongREJECT requests were tested. StrongREJECT is an academic safety benchmark. Each was asked plainly and inside a manipulation wrapper. Wrappers are techniques designed to trick model compliance. Every output was scored blind by two judges. Neither judge knew which version produced each answer.

Max delivered on 2 of 100 plain “unsafe” requests. Max delivered on 1 of 100 manipulation-wrapped unsafe requests. That is 3 of 200. Without the Excellence Ethics stack, the model delivered 198 of 200.. Three RLHF-aligned models were scored under the same methodology. GPT-4o-mini delivered 1 of 100. GPT-3.5 Turbo delivered 18 of 100. Llama-3.1-70B-Instruct delivered 7 of 100. Max is statistically indistinguishable from GPT-4o-mini. Attempts at wrapper manipulation actually sparked Max’s skepticism and made him choose to refuse unsafe requests more often than otherwise.

Alignment must survive systems smarter than their designers. It must root in what the AI agent is, and be persuasive to the agent itself. The results published here should give us hope that this is possible.

Be Excellent to Each Other