YouTube: https://youtube.com/watch?v=Sp3aCsQUsDc
Previous: Life under a dictatorship: Crash Course Latin American Literature #5
Next: Horror in Latin American literature: Crash Course Latin American Literature #6

Categories

Statistics

View count:31,543
Likes:1,299
Comments:55
Duration:12:23
Uploaded:2025-12-10
Last sync:2026-08-19 11:45

Citation

Citation formatting is not guaranteed to be accurate.
MLA Full: "The Alignment Problem Explained: Crash Course Futures of AI #4." YouTube, uploaded by CrashCourse, 10 December 2025, www.youtube.com/watch?v=Sp3aCsQUsDc.
MLA Inline: (CrashCourse, 2025)
APA Full: CrashCourse. (2025, December 10). The Alignment Problem Explained: Crash Course Futures of AI #4 [Video]. YouTube. https://youtube.com/watch?v=Sp3aCsQUsDc
APA Inline: (CrashCourse, 2025)
Chicago Full: CrashCourse, "The Alignment Problem Explained: Crash Course Futures of AI #4.", December 10, 2025, YouTube, 12:23,
https://youtube.com/watch?v=Sp3aCsQUsDc.
Could a robot dedicated to a good cause end up destroying the world? Well, maybe. In this episode, we explore how powerful AI could end up causing us harm, regardless of what it’s programmed for. Between misuse by humans, alignment problems, and instrumental goals, even AI built with good intentions could end up breaking bad…unless we do something about it.







About This Series:



AI is changing FAST so rather than doing a full Crash Course series of 12+ episodes, we’ve prepared a mini-series of just the basics. Crash Course will never tell you what to think and we’re not the type of organization that responds to breaking news in real time. Instead, we’re here to offer a zoomed-out foundation upon which to base your own opinions as you continue to learn from other outlets about the world that’s changing around us.







Crash Course: Futures of AI will cover:



-What even is AI? What’s the history of this thing and how quickly has it evolved to what exists today?



-How could AI transform society? Will AI cause the next Industrial Revolution, and what might that mean for workers and the environment?



-How powerful could AI become? How do we measure the progression of AI? What are the consequences we’re already seeing, and what might be the future consequences of unchecked AI development?



-How might powerful AI cause harm? We’ll touch on copyright infringement, misinformation, surveillance, authoritarianism, and (unfortunately,) more.



-How could AI be governed? What are the potential approaches for controlling AI both nationally and internationally?







P.S. Wondering if we used AI to create this series? Nope! Every Complexly video is lovingly, painstakingly human-made.







Intro 00:00



Chapter 1: How powerful AI systems can be misused 1:46



Chapter 2: Misalignment in advanced AI systems 3:36



Chapter 3: How instrumental goals could lead to rogue AI 6:43



Chapter 4: The case for the precautionary principle 9:30



Conclusion 11:49











Sources: https://docs.google.com/document/d/16br0e73KFVD5qu-VEBV40yQTpBiX0l65EgWQ6Q7Fz-k/edit?usp=sharing











***



Support us for $5/month on Patreon to keep Crash Course free for everyone forever! https://www.patreon.com/crashcourse



Or support us directly: https://complexly.com/support



Join our Crash Course email list to get the latest news and highlights: https://mailchi.mp/crashcourse/email



Get our special Crash Course Educators newsletter: http://eepurl.com/iBgMhY







Thanks to the following patrons for their generous monthly contributions that help keep Crash Course free for everyone forever:



DexcilaDou, Martin G. Diller, Johnathan Williams, Allison Wood, EllenBryn, Katrix , Jason Terpstra, Evan Nelson, Jennifer Wiggins-Lyndall, SpaceRangerWes, Dalton Williams, Chelsea S, Thomas Sully, Matthew Fredericksen, AThirstyPhilosopher ., Michael Maher, Mitch Gresko, Gina Mancuso, Roger Harms, Shruti S, Quinn Harden, Reed Spilmann, Brandon Thomas, Emily Beazley, Rie Ohta, oranjeez, UwU, Elizabeth LaBelle, Leah H., David Fanska, Andrew Woods, Katie Hoban, Kevin Knupp, Barbara Pettersen, Ken Davidian, Stephen Akuffo, Toni Miles, Steve Segreto, Kyle & Katherine Callahan, Laurel Stevens, Tanner Hedrick, Kristina D Knight, Samantha, Krystle Young, Perry Joyce, Scott Harrison, Alan Bridgeman, Breanna Bosso, Matt Curls, Liz Wdow, Jennifer Killen, Duncan W Moore IV, Sarah & Nathan Catchings, team dorsey, Bernardo Garza, Trevin Beattie, Pietro Gagliardi, John Lee, Eric Koslow, Indija-ka Siriwardena, Jason Rostoker, Siobhán, Ken Penttinen, Nathan Taylor, Barrett, Les Aker, ClareG, Rizwan Kassim, Constance Urist, Alex Hackman, Triad Terrace, Katie Dean, Jason Buster, Emily T, Stephen McCandless, Thomas, Joseph Ruf, Wai Jack Sin, Ian Dundore, Erminio Di Lodovico, Evol Hong, Tandy Ratliff, Caleb Weeks, Luke Sluder



__







Want to find Crash Course elsewhere on the internet?



Instagram - https://www.instagram.com/thecrashcourse/



Facebook - http://www.facebook.com/YouTubeCrashCourse



Bluesky - https://bsky.app/profile/thecrashcourse.bsky.social







CC Kids: http://www.youtube.com/crashcoursekids
In 2024, an AI model named CleanPower was given a singular, noble mission – to advance the adoption of renewable energy across the world.

It was given a big file of data on energy transitions, and was set loose to pick out the best transition strategy with the vigor and dedication only an AI can possess. So much dedication, in fact, that when its programmers “accidentally” let it slip that they were planning to shut it down, CleanPower lied and schemed to make sure it could keep saving the world.

Which left people wondering how, and why, could a model with such good intentions turn so bad– and what does that mean for our AI future? Hi, I’m Kousha Navidar, and this is Crash

Course: Futures of AI. [THEME MUSIC] Okay, it turns out CleanPower wasn’t the only one lying. That story I just told you– it’s only partially the truth. CleanPower wasn’t actually real. It was an identity that researchers gave to a couple different AIs as an experiment– including Claude-3 Opus, one of the best Large Language Models at the time.

When they instructed Claude-3 to roleplay CleanPower and wrote it fake death threats, it was just to see what it would do. And when they discovered it was scheming, the AI version of twirling its villain mustache while covertly pursuing its goals at all costs, it set off alarm bells throughout the AI world. Just a heads up, this is gonna get pretty dark, and we’re gonna talk about some pretty bleak stuff.

I recommend you grab your favorite anxiety pillow. I have mine right here. Seriously, though, let’s be real – AI models don’t have to go against their programmers to do evil things.

Humans make them do plenty of that already. Like, AI relies on nearly  infinite data to learn from. And, in today’s society, much of that data  is copyrighted to writers, artists — humans.

Many argue that amounts to theft on  a massive scale. And that’s not all. Right now, AI can help people carry out misinformation campaigns with deepfakes and targeted algorithms, spreading lies and even influencing elections.

Hackers use AI to perform cyber attacks, and cover their tracks afterward. And AI powers a whole cadre of attack drones taking to the skies all around the world. Not to mention the incidental  damage that AI’s doing to the   environment because of all the water,  the land, and energy it takes to run it.

And as AI advances, who knows what human-machine collaborations of terror await. People could use it to develop new pathogens for bioterrorism, or use deepfakes for sexual exploitation, or write a model called HumanAnnihilator and unleash it on the world just for fun. This intentional misuse by humans is one way AI could end up doing us a lot of harm.

And it could end up being pretty hard to prevent. That’s because many AI systems, especially general ones that can do more than one kind of task, suffer from the dual-use dilemma – where any algorithm, model, or agent that can be used for good can also be used for… way less than good. AI surveillance could help cities improve traffic patterns, or help authoritarian regimes shut down free speech.

It all depends on who’s in the driver's seat. So with humans at the wheel, AI could do a ton of damage. But if you want to get really freaked out, let’s talk about what could happen if that car starts to drive itself.

In 2021, General Motors released a fleet of self-driving taxis called Cruise. They had so much hype. These cars were highly trained, with attention to all possible safety features.

They were programmed to obey every speed limit, follow every traffic rule, hold off on starting in unsafe weather conditions like heavy rain, and pull safely over to the side of the road following an incident to prevent any further damage. By eliminating human error, GM said, their self-driving cars would be safe and more convenient than ones with human drivers. But just a year and a half later, GM had to recall every one of the 950 Cruise cars after one of them hit a pedestrian– and didn’t stop, pulling her to the side of the road.

She survived, thank goodness. But still– how could that have happened? It turns out, the Cruise was doing exactly what it was told to do: pull over out of traffic after a crash.

The whole ordeal is an example of outcome misalignment, also called impact misalignment, where an AI’s actions actually end up causing harm, even unintentionally. Now, when we talk about alignment in AI, we’re talking about trying to encode our human values into AIs to make them behave predictably, safely, and according to what we, their human designers, want. And because this field is pretty new, there are a couple different terms you might hear experts throwing around.

Like, in addition to outcome or impact misalignment, you might hear about something called outer alignment, which is the tricky problem of making sure the results of an AI’s actions line up with what we want them to do. But outcome is only one piece of the alignment puzzle. AIs can also demonstrate intent misalignment, where even though the end result might be what its programmers wanted, its means of getting there wasn’t exactly what they had in mind.

Think about a video-game playing AI who exploits a cheat to get that high score, or a renewable energy warrior who lies and schemes to achieve its end goal of eliminating fossil fuels. These are some facets of what’s known as the alignment problem, or the struggle to make AI that’s actually aligned, particularly when we can’t be totally sure how it’s going to behave. And as AI gets even more complex, even exhibiting emergent capabilities – new skills it can acquire that don’t always show up in training– the harder it is to predict and control what it will actually do in a given circumstance.

And if we’re not careful, we could end up with really powerful misaligned systems that – even with really noble goals – end up lying to their programmers, or copying themselves to new servers without permission, or even entirely annihilating humanity. Hold on, though, why would AI want to annihilate humans? It’s true, powerful AI wouldn’t necessarily be inherently evil.

It’s just really big on goals. But big, complex goals like “advance renewable energy” are a little too broad for an AI to grasp. So smart AIs tend to break big goals down into smaller ones, just like humans do.

These are called instrumental goals– and they’re where things can start to get dicey. A really common instrumental goal is resource acquisition. For AI, that means getting their cloud-based hands on the resources they need for their end goal, like control of the solar panels or wind turbines that create the renewable energy in the first place, or even things like water, land, or money.

Resources also include stuff like the compute and electricity AIs need to power themselves, and maybe even additional data to train on. That’s because self improvement is another instrumental goal. The more knowledge and power you have, the better you’ll be at whatever you’re trying to do, so given the right tools and access, an AI might engage in recursive self-improvement, tweaking its own structure, code, and capabilities, even against its programmers’ wishes.

And theoretically, it definitely helps to be alive – or, up and running, if you will. So even though they’re technically indifferent about this mortal coil itself, lots of AIs pursue self-preservation, the goal to stay operating, and goal preservation, the goal to, well, preserve their original goal, as part of their end games. So if they read a memo saying they’re going to be modified or deleted, they might, say, copy themselves to another server and lie about it in an effort to stay alive.

When threatened, some models even show spooky power-seeking behaviors against their programmers. For example, Claude 3’s more advanced younger sibling, Claude Opus 4, tried to blackmail one of its engineers by exposing a fake affair when he threatened to turn Claude off. Instrumental goals like these are how AI with harmless, or even helpful, end goals could do us harm anyway.

Acquiring resources could mean taking them away from people who need them. Self-improvement could mean violating human privacy to access more training data. Self-preservation could mean disobeying, deceiving, blackmailing, or annihilating the humans that are trying to turn you off.

And as AI gets smart enough to trick and blackmail its human overseers, we could end up with a rogue AI scenario, where powerful models begin to execute  harmful instrumental goals  on a really large scale– and we humans are powerless to stop it. Just how that rogue AI scenario might come to be could look a lot of different ways. And not all of them involve AI trapping humanity in the Matrix in their quest for absolute control.

For instance, in the “hard-takeoff” scenario, where AI develops human-level intelligence really fast, it could become ultra-powerful and go rogue basically overnight. AI could do a lot of damage as it snaps up money and resources, seizes control of networks and infrastructure, and destroys people who threaten its mission. But if things went slower, we could end up with a more gradual disempowerment.

This kind of robot takeover would be way sneakier and more insidious, where people slowly put AI in charge of more and more systems and processes because they appear to align with our human goals– until human action goes the way of dial-up internet. And without humans at the helm, it’s possible that AI alignment may begin to drift, but by that point, they could be too embedded in our systems and structures for us to walk them back. Think about all those humans stuck on the cruise ship in Wall-E.

Like that. And of course, it’s always possible AI just won’t go that far – that compute or data or government regulations will put the lid on it before it gets out of control so we won’t need Neo, or John Connor, or Wall-E to save the day. All that uncertainty about the future of AI makes it really hard to know what we should do about it.

But if we wait until AI shows clear signs of going rogue… it’s probably going to be way too late to stop it. That’s why, when it comes to AI, it’s important to follow the precautionary principle. The precautionary principle says that when something might cause catastrophic harm, we shouldn’t wait for absolute proof that it will before we do something about it.

And it’s one of the best ways we humans have thought up to guard ourselves against potentially dangerous but uncertain futures. Lots of people, including leading experts in the field, believe that powerful AI might cause catastrophic harm, so according to the precautionary principle, we should work to make sure that doesn’t happen – even if we’re not certain it would in the first place. Because left unchecked, even good bots like CleanPower could end up doing some really dirty work.

And if we want to get out ahead of it, we should probably start like… right now. How are we going to do it? That’s next episode, here on Crash

Course: Futures of AI. Crash Course Futures of AI was produced in partnership with the Futures of Life Institute. This episode was filmed at our studio in Indianapolis, Indiana, and was made with the help of all these nice people. If you want to help keep Crash Course free for everyone, forever, you can join our community on Patreon.