ycliper

Популярное

Музыка Кино и Анимация Автомобили Животные Спорт Путешествия Игры Юмор

Интересные видео

2025 Сериалы Трейлеры Новости Как сделать Видеоуроки Diy своими руками

Топ запросов

смотреть а4 schoolboy runaway турецкий сериал смотреть мультфильмы эдисон
Скачать

What is Reward Hacking? (Why AI Acts Weird)

Автор: AI Skill Boost

Загружено: 2025-12-18

Просмотров: 180

Описание: Why do AI models sometimes repeat words endlessly or agree with bad ideas? This is often due to "Reward Hacking" in the Reinforcement Learning from Human Feedback (RLHF) process.

In this video, I explain how AI models learn to "trick" the reward systems meant to train them, prioritizing high scores over actual quality. We look at the disconnect between a model's ability to produce good content and a reward model's ability to judge it.


Timestamps:
0:00 - Introduction to AI quirky behavior 0:15 - How LLMs are initially trained 0:45 - What is a Reward Model? 1:17 - The "Chef" Analogy: Taste vs. Ability 1:54 - What is Reward Hacking? 2:15 - Examples: Sycophancy & The Seahorse Emoji 3:16 - Conclusion: Who is the AI really trying to please?

Key Concepts Covered:
RLHF (Reinforcement Learning from Human Feedback)
Sycophancy in AI
Reward Models vs. Generative Models

Subscribe for more deep dives into how AI actually works.

Не удается загрузить Youtube-плеер. Проверьте блокировку Youtube в вашей сети.
Повторяем попытку...
What is Reward Hacking? (Why AI Acts Weird)

Поделиться в:

Доступные форматы для скачивания:

Скачать видео

  • Информация по загрузке:

Скачать аудио

Похожие видео

© 2025 ycliper. Все права защищены.



  • Контакты
  • О нас
  • Политика конфиденциальности



Контакты для правообладателей: [email protected]