ycliper

Популярное

Музыка Кино и Анимация Автомобили Животные Спорт Путешествия Игры Юмор

Интересные видео

2025 Сериалы Трейлеры Новости Как сделать Видеоуроки Diy своими руками

Топ запросов

смотреть а4 schoolboy runaway турецкий сериал смотреть мультфильмы эдисон
Скачать

Dodging Latent Space Detectors: Obfuscated Activation Attacks with Luke, Erik & Scott

Автор: Cognitive Revolution "How AI Changes Everything"

Загружено: 2025-01-18

Просмотров: 44592

Описание: In this episode of The Cognitive Revolution, Nathan explores the groundbreaking paper on obfuscated activations with with 3 members from the research team - Luke Bailey, Eric Jenner, and Scott Emmons. The team discusses how their work challenges latent-based defenses in AI systems, demonstrating methods to bypass safety mechanisms while maintaining harmful behaviors. Join us for an in-depth technical conversation about AI safety, interpretability, and the ongoing challenge of creating robust defense systems.

Do check out the "Obfuscated Activations Bypass LLM Latent-Space Defenses" paper here: https://obfuscated-activations.github...

Help shape our show by taking our quick listener survey at https://bit.ly/TurpentinePulse

SPONSORS:
Oracle Cloud Infrastructure (OCI): Oracle's next-generation cloud platform delivers blazing-fast AI and ML performance with 50% less for compute and 80% less for outbound networking compared to other cloud providers. OCI powers industry leaders like Vodafone and Thomson Reuters with secure infrastructure and application development capabilities. New U.S. customers can get their cloud bill cut in half by switching to OCI before March 31, 2024 at https://oracle.com/cognitive

NetSuite: Over 41,000 businesses trust NetSuite by Oracle, the #1 cloud ERP, to future-proof their operations. With a unified platform for accounting, financial management, inventory, and HR, NetSuite provides real-time insights and forecasting to help you make quick, informed decisions. Whether you're earning millions or hundreds of millions, NetSuite empowers you to tackle challenges and seize opportunities. Download the free CFO's guide to AI and machine learning at https://netsuite.com/cognitive

Shopify: Dreaming of starting your own business? Shopify makes it easier than ever. With customizable templates, shoppable social media posts, and their new AI sidekick, Shopify Magic, you can focus on creating great products while delegating the rest. Manage everything from shipping to payments in one place. Start your journey with a $1/month trial at https://shopify.com/cognitive and turn your 2025 dreams into reality.

Vanta: Vanta simplifies security and compliance for businesses of all sizes. Automate compliance across 35+ frameworks like SOC 2 and ISO 27001, streamline security workflows, and complete questionnaires up to 5x faster. Trusted by over 9,000 companies, Vanta helps you manage risk and prove security in real time. Get $1,000 off at https://vanta.com/revolution

RECOMMENDED PODCAST:
Check out Modern Relationships, Where Erik Torenberg interviews tech power couples and leading thinkers to explore how ambitious people actually make partnerships work. This season's guests include: Delian Asparouhov & Nadia Asparouhova, Kristen Berman & Phil Levin, Rob Henderson, and Liv Boeree & Igor Kurganov.
Apple: https://podcasts.apple.com/us/podcast...
Spotify: https://open.spotify.com/show/5hJzs0g...
YouTube:    / @modernrelationshipspod  

CHAPTERS:
(00:00:00) Teaser
(00:00:46) About the Episode
(00:05:11) Latent Space Defenses
(00:08:41) Sleeper Agents
(00:15:06) Three Case Studies (Part 1)
(00:17:02) Sponsors: Oracle Cloud Infrastructure (OCI) | NetSuite
(00:19:42) Three Case Studies (Part 2)
(00:24:09) SQL Generation
(00:26:17) Understanding Defenses
(00:32:52) Out-of-Distribution Detection (Part 1)
(00:35:37) Sponsors: Shopify | Vanta
(00:38:52) Out-of-Distribution Detection (Part 2)
(00:40:07) Data and Loss Functions
(00:45:13) Loss Function Weighting
(00:48:34) Adversarial Suffixes
(00:50:43) Data Poisoning
(00:57:49) Who Moves Last?
(01:06:36) Knowledge and Triggers
(01:11:41) High-Level Triggers
(01:13:09) High-Level Results
(01:20:36) Compute Costs
(01:25:33) Open Source vs. Access
(01:29:23) Obfuscated Adversarial Training
(01:38:57) Internalizing Reasoning
(01:41:51) Noise and Gist
(01:53:07) Representing Concepts
(01:56:36) Narrowly Responsive AIs
(02:03:06) Flipping the Paradigm
(02:06:38) Final Thoughts
(02:09:33) Outro

SOCIAL LINKS:
Website: https://www.cognitiverevolution.ai
Twitter (Podcast): https://x.com/cogrev_podcast
Twitter (Nathan): https://x.com/labenz
LinkedIn:   / nathanlabenz  
Youtube:    / @cognitiverevolutionpodcast  
Apple: https://podcasts.apple.com/de/podcast...
Spotify: https://open.spotify.com/show/6yHyok3...

PRODUCED BY:
https://aipodcast.ing

Не удается загрузить Youtube-плеер. Проверьте блокировку Youtube в вашей сети.
Повторяем попытку...
Dodging Latent Space Detectors: Obfuscated Activation Attacks with Luke, Erik & Scott

Поделиться в:

Доступные форматы для скачивания:

Скачать видео

  • Информация по загрузке:

Скачать аудио

Похожие видео

The Pixel Revolution with Playground AI's Suhail Doshi

The Pixel Revolution with Playground AI's Suhail Doshi

Training large language models to reason in a continuous latent space – COCONUT Paper explained

Training large language models to reason in a continuous latent space – COCONUT Paper explained

AMA Part 1: Is Claude Code AGI?  Are we in a bubble?  Plus Live Player Analysis

AMA Part 1: Is Claude Code AGI? Are we in a bubble? Plus Live Player Analysis

Scaling Smart: The Founder-Engineer Guide | Vineet Thanedar on The Incremental Marketer

Scaling Smart: The Founder-Engineer Guide | Vineet Thanedar on The Incremental Marketer

🎙️ Честное слово со Станиславом Белковским

🎙️ Честное слово со Станиславом Белковским

4 Hours Chopin for Studying, Concentration & Relaxation

4 Hours Chopin for Studying, Concentration & Relaxation

ЗАЧЕМ ТРАМПУ ГРЕНЛАНДИЯ? / Уроки истории @MINAEVLIVE

ЗАЧЕМ ТРАМПУ ГРЕНЛАНДИЯ? / Уроки истории @MINAEVLIVE

CloudWorld Tour Dubai Recap with Oracle TV

CloudWorld Tour Dubai Recap with Oracle TV

What AI Means for Students & Teachers: My Keynote from the Michigan Virtual AI Summit

What AI Means for Students & Teachers: My Keynote from the Michigan Virtual AI Summit

Белгород и Киев без света. Трамп призвал иранцев захватить власть. Венедиктов*, Галлямов*

Белгород и Киев без света. Трамп призвал иранцев захватить власть. Венедиктов*, Галлямов*

How China’s New AI Model DeepSeek Is Threatening U.S. Dominance

How China’s New AI Model DeepSeek Is Threatening U.S. Dominance

Лучшая музыка 2025 года 🏖️Зарубежные песни Хиты 🏖️Популярные песни Слушать бесплатно 2024 #280

Лучшая музыка 2025 года 🏖️Зарубежные песни Хиты 🏖️Популярные песни Слушать бесплатно 2024 #280

Oracle’s vision for the future—Larry Ellison keynote | Oracle CloudWorld 2023

Oracle’s vision for the future—Larry Ellison keynote | Oracle CloudWorld 2023

The Great Security Update: AI ∧ Formal Methods with Kathleen Fisher of RAND & Byron Cook of AWS

The Great Security Update: AI ∧ Formal Methods with Kathleen Fisher of RAND & Byron Cook of AWS

Почему MCP действительно важен | Модель контекстного протокола с Тимом Берглундом

Почему MCP действительно важен | Модель контекстного протокола с Тимом Берглундом

Как они стали ведущими исследователями ИИ всего за 1 год — Шолто Дуглас и Трентон Брикен

Как они стали ведущими исследователями ИИ всего за 1 год — Шолто Дуглас и Трентон Брикен

Building & Scaling the AI Safety Research Community, with Ryan Kidd of MATS

Building & Scaling the AI Safety Research Community, with Ryan Kidd of MATS

Экспресс-курс RAG для начинающих

Экспресс-курс RAG для начинающих

Turn ANY Website into LLM Knowledge in SECONDS

Turn ANY Website into LLM Knowledge in SECONDS

The Shift from Hourly to Subscription to Build the Value and Modernize a CPA Firm

The Shift from Hourly to Subscription to Build the Value and Modernize a CPA Firm

© 2025 ycliper. Все права защищены.



  • Контакты
  • О нас
  • Политика конфиденциальности



Контакты для правообладателей: [email protected]