Is RLHF good for AI safety?
2
Sep 27
Yes
No
See results
This question is managed and resolved by Manifold.
Get
1,000 to start trading!
People are also trading
Related questions
Does RL harm off-target capabilities? Is that temporary?
50% chance
Will prioritizing corrigible AI produce safe results?
45% chance
What grade will OpenAI get for existential safety in the FLI's AI Safety Index Winter 2026?
How will the structural similarity of the three frontier AI safety frameworks (Anthropic RSP, OpenAI Preparedness, GDM F
What grade will Z.ai get for existential safety in the FLI's AI Safety Index Winter 2026?
What grade will Anthropic get for existential safety in the FLI's AI Safety Index Winter 2026?
Will something AI-related be an actual infohazard?
68% chance
Will AI be considered safe in 2030? (resolves to poll)
72% chance
In 2030, will we think FLI's 6 month pause open letter helped or harmed our AI x-risk chances?
Will "We should try to automate AI safety work asap" make the top fifty posts in LessWrong's 2025 Annual Review?
14% chance