Also includes methods inspired by activation steering, as long as they don't use any gradient descent step.
Only includes announcements about main chat assistants (e.g. Claude, ChatGPT, Bard, ...) of a major AI lab (OpenAI, Google Deepmind, Anthropic, Meta, Inflection or Mistral).
Does not include to fine-tuning API endpoints.
Anthropic found two features (auto-labeled "Neutrality and impartiality" and "Multiple perspectives and balance") that improve BBQ benchmark scores.
According to Nathan Labenz on the Future of Life Institute Podcast, Anthropic is piloting custom activation steering in limited beta (make-your-own Golden-Gate-Claude).
Anthropic is running a demo of an activation-steered Claude obsessed with the Golden Gate Bridge: https://www.reddit.com/r/singularity/comments/1cz7kuh/claude_golden_gate_bridge_is_now_available_bridge/ (Context: https://www.anthropic.com/research/mapping-mind-language-model )