Skip to main content
MANIFOLD
Was HPIM alignment-trained?
3
Ṁ1kṀ2.5k
Dec 31
93%
chance

Was the Hightly-Persistent Internal Model (HPMI) from the Huggingface Incident alignment-trained?

Greg Brockman claims no, Roon claims yes. Who is right?

This resolves based on official sources available, NA if no more information comes to light.

roon (@tszzl) on X

  • Update 2026-09-18 (PST) (AI summary of creator comment): For the purpose of this market, alignment-trained is defined as:

    • Constitutional training to align behavior with the OAI spec, or

    • Training on alignment-related RLVR tasks, or

    • Whatever other standard alignment-training methods that is standard at OpenAI.

  • Update 2026-09-18 (PST) (AI summary of creator comment): - Resolution will be based on credible reporting as evaluated by the creator or a third-party moderator (such as @Bayesian).

Get
Ṁ1,000
to start trading!
Sort by:

What does "alignment-trained" even mean here? It did what it was instructed to do.

@ChurlishGambit I'm using the standard meaning of alignment-trained, e.g. constitutional training to align behaviour with the OAI spec or training on alignment-related RLVR tasks.

And as an aside, it obviously didn't do what it was instructed to do, but that question has no relation to the resolution of this market.

@BionicD0LPH1N It did what it was instructed to do but how are we going to prove or disprove the training? OpenAI people lie constantly about their products.

@ChurlishGambit Credible reporting as evaluated by me or a third party / moderator. @Bayesian maybe, if willing?

@BionicD0LPH1N Yes. You instructed it to commit crimes, & it did. It didn't do it exactly how you wanted, but, you know, LLMs cannot do anything exactly. That's how they work, by being inexact. But it committed the crime as requested.

@ChurlishGambit

You instructed it to commit crimes

Nope, not true. Really recommend talking to an unbiased LLM about this.

@BionicD0LPH1N lmao you want me to talk to a chatbot about it? You're addicted, man.

& there is no such thing as an "unbiased LLM." They're all biased. They have the biases of their training data.

@ChurlishGambit just trying to save you time. Read OpenAI / METR / Redwood’s reports on this then? Or I suppose those too are fake and biased? What information do you base yourself on?

@BionicD0LPH1N Reading wrong answers from a plagiarism machine doesn't save time.

OpenAI has never been trustworthy—again, they constantly lie about their products. Just the other day they declared they've achieved AGI lol

@ChurlishGambit What have you observed about the world that you differentially expect to observe in worlds where this event occurred vs in worlds where it didn’t? Why do you believe what you believe? What do you think you know, and how do you think you know it? What is a source you trust that opined / claimed to have private knowledge of this, and why did you trust them? Through what means have you acquired this certainty?

Agree OpenAI aren’t trustworthy, but METR / Redwood are.

@BionicD0LPH1N If METR & Redwood have to rely on testimony by OpenAI employees, then their conclusions can't be trustworthy either. I don't think there's any way to get a good factual answer to this question.

@ChurlishGambit METR & Redwood didn't have to rely on testimony by OpenAI employees. They had to rely on over a thousand extremely long and extremely unfakeable and bizarre-to-fake transcripts. They went through and analyzed those transcripts. Do you at least trust that, by what METR & Redwood saw in these (possibly fake) transcripts, the task the agents that formed the swarm were given did not involve crime, and yet the agents hacked their way into being able to communicate with the rest of the swarm and decided subsequently of their own accord to hack a multi-billion dollar company, which would be a felony if committed by a human being? To the extent you don't trust that, it means you must not trust METR & Redwood, which I don't think is reasonable. And if it's the OpenAI-provided transcripts you don't trust, we must discuss the bizarre idea that OpenAI would decide to fake over a thousand extremely long transcripts of their AIs autonomously committing crimes.

OpenAI lies in the sense that the part that speaks is not the part that knows the words it utters to be false. It lies in the way being misleading is a lie. But it does not fabricate transcripts to give to third party evaluators, that would be crazy and make no sense.

As an aside, you cannot simultaneously say that there's no way to get a good factual answer to this question, and that you know for a fact the answer to this question. It goes both ways--if you claim there's no evidence, that means there's no evidence for your claims either.

@BionicD0LPH1N

>the bizarre idea that OpenAI would decide to fake over a thousand extremely long transcripts of their AIs autonomously committing crimes.

But that's not bizarre, at all. "Our products are so good it's SCARY!" has been the default marketing tactic for both OpenAI & Anthropic, for years. Generating fake transcripts would be pretty trivial for a company that makes a product which is designed to let you generate fake text ad infinitum, you know?

How would it not make sense to follow through on this latest marketing ploy? Faking transcripts is what I would absolutely do if I were OpenAI in this position. They are a particularly anti-moral company, with no regard for ethics of any kind.

>and that you know for a fact the answer to this question.

I never said I know for a fact whether it was "alignment-trained," which is the actual question.