DeepStack: Artificial Intelligence Learns to Bluff and Beats Poker Pros

Article cover

In chess, both players see all the pieces, and in Go, every stone’s position is known. Poker is different. A player does not see their opponent’s cards and can never be sure if a bet represents a strong hand or just a well-timed bluff. Additionally, the outcome of a hand doesn’t always reveal who made the correct decision. A player can play perfectly and still lose or make a major mistake and win the entire pot.

This is why poker was long considered one of the toughest tests for artificial intelligence. A computer cannot merely crunch probabilities; it has to deal with incomplete information, estimate the opponent's strategy, and remain unpredictable. In 2017, researchers introduced a system that showed AI can tackle this challenge. Named DeepStack, it became the first program to decisively beat professional players in heads-up no-limit Texas hold’em.

DeepStack was an AI research system developed specifically for heads-up no-limit Texas hold’em, a game between two players without predetermined bet limits. Researchers from the University of Alberta, Charles University in Prague, and the Czech Technical University collaborated on its development. Their goal was to create a strategy for a game where not all information is known, opponents actively deceive each other, and each bet can have numerous meanings. While DeepStack wasn't a universal AI, it excelled in the task it was designed for, surpassing professional players.

Why Is Poker So Challenging for Computers?

At first glance, Texas hold’em might seem simpler than chess. It uses only 52 cards, and each decision offers a few basic actions: check, bet, call, raise, or fold. The true complexity, however, comes from uncertainty. Players can't see their opponent's cards, so they must consider entire sets of possible hands, known in poker as ranges.

If an opponent raises before the flop, they might have a high pair, strong ace, suited connectors, or even a weaker hand. Subsequent decisions can make certain hands more or less likely, gradually narrowing the possible range. Players must ask more than just, “What does my opponent have?” They need to consider what cards the opponent could hold, how often they might have them, and what their previous actions reveal about their entire range. This process is reciprocal; every bet conveys information but can also be a bluff.

How DeepStack Learned to Bluff

A player who bets only with very strong hands can be easily beaten. Once they commit to the pot, the opponent can fold weaker cards without much fear. Similarly, a player who bluffs too often is also vulnerable; the opponent will start calling bets more frequently. A strong poker strategy requires balance, using the same action with strong hands, medium-strength hands, and bluffs.

DeepStack didn’t seek simple answers like “always bet in this situation.” Instead, it developed probabilistic strategies, sometimes checking, using small bets, or opting for large sizing in the same situation to keep its strategy unreadable and unexploitable by opponents. This approach naturally included bluffs. DeepStack didn’t bluff due to courage, fear, or a desire to psychologically intimidate. For him, bluffing was a mathematical part of a balanced strategy.

Heads-up no-limit hold’em contains roughly 10¹⁶⁰ decision points. Every player can hold different cards, thousands of board combinations can appear, and varying bet sizes create more continuations. This decision tree is too vast to fully process, even with significant computational power.

Older poker programs used extensive simplifications, grouping similar situations, and allowing only predefined bet sizes instead of all possibilities. If a player chose a different bet than the system expected, the program had to match it to the nearest known option, which could introduce weaknesses. DeepStack chose a different approach. It recalculated strategy during play instead of pre-solving every potential situation.

Neural Network as Poker Intuition

The system was grounded in a principle known as continual re-solving, constantly solving the current situation. This approach is akin to automobile navigation, which doesn’t need separate plans for every potential traffic jam or wrong turn. Knowing the current location, destination, and surroundings, it recalculates the next segment of the journey as circumstances change.

DeepStack worked similarly. At each decision, it evaluated the pot size, community cards, and possible ranges for both players, then solved the most relevant part of the game from there. After the opponent acted, it repeated the process. If analyzing further decisions became too complex, DeepStack used a neural network.

This network estimated the value of the current situation based on the pot size, board, and both players' ranges, allowing DeepStack to forego calculating every possible continuation through to the river, instead replacing further predictions with a sufficiently accurate estimate. This strategy combined various approaches: game theory for balanced strategy, search algorithms for immediate decisions, and neural networks for predicting distant outcomes.

DeepStack didn't need to analyze millions of hands from top players or mimic human decisions. Researchers generated huge amounts of random situations with varying boards, pot sizes, and ranges, solving them with game theory algorithms and using the results as training data for the neural network.

DeepStack learned not how a successful pro would play a particular hand, but what a strategically advantageous solution looked like. While training required significant computational resources, DeepStack could function during actual play on a single NVIDIA GeForce GTX 1080 graphics card, averaging three seconds per decision.

The Battle Against Poker Pros

The real test came at the end of 2016, when 33 poker players from 17 countries faced DeepStack, collectively playing 44,852 hands. Each participant was initially to play 3,000 hands, and 11 players completed the full sample. Adjusted results showed DeepStack profiting against all 11, achieving a statistically significant victory individually against ten of them.

Even against the most successful participant, the system was estimated to be profitable, though not conclusively. Overall, DeepStack outperformed the group by over four standard deviations from zero, which researchers described as a highly significant result.

In poker, simply comparing chip counts isn’t enough. A weaker player might get better cards over thousands of hands, win more all-ins, and finish profitably. Researchers used the AIVAT method to reduce the influence of luck, utilizing known probabilities and expected values to separate decision quality from short-term variance. AIVAT reduced the standard deviation of outcomes by roughly 85 percent, leaving DeepStack's victory impressive even with these adjustments.

Did DeepStack Solve Poker?

Despite defeating professionals, this didn’t mean DeepStack mathematically solved heads-up no-limit hold’em. The game is too expansive to determine exactly how close its strategy was to perfect balance. Although researchers sought weaknesses with special tests, they didn’t find any easy way to consistently beat the system.

DeepStack could only play heads-up no-limit hold’em, handling neither 6-max cash games nor full-ring or tournaments with changing blinds. It wasn’t equipped to deal with ICM, payout structure, or the dynamics of multiple players, nor switch automatically to Pot-Limit Omaha or mixed games. Matches were online, without observing physical tells, facial expressions, or opponent anxiety. The goal wasn’t to develop a psychological profile but to use a robust strategy that could not be easily exploited.

Despite this, DeepStack marked the beginning of a new era. 2017 was a pivotal year for poker AI. Around the same time, Carnegie Mellon University’s Libratus system played 120,000 hands against four elite heads-up professionals, also achieving a convincing victory.

Another major step came in 2019 when the system Pluribus defeated professionals in 6-max no-limit hold’em. Progress was rapid: by 2015, AI had practically solved heads-up limit hold’em, two years later it was beating pros in no-limit heads-up, and by 2019, it managed to win at tables with multiple players.

Before the solver era, players learned mainly from experience, discussion, training videos, and analyzing successful professionals. They knew balanced ranges and sufficient bluffs were necessary, but exact frequencies were often guessed. Modern solvers offered far more detailed insights, showing why one combination suited a bluff while a similar one didn’t, why a small bet made sense on one board and large sizing on another, or why certain hands were mixed between checks and bets.

The Dark Side of Poker AI

Developing strong poker AI raises an unsettling question: what if someone uses this technology in online play? Players don’t have to grant programs full control over accounts. Simply getting real-time advice from external software is enough. This kind of help is known as real-time assistance (RTA). Using solvers post-game is standard for study, but live assistance during play is prohibited on online poker platforms.

Advanced systems needn’t appear mechanical, capable of blending strategies, applying randomness, and behaving naturally. Poker sites thus analyze decision timing, account behaviors, devices used, possible links between players and strategic similarities. DeepStack was an academic project, not a cheating tool, yet demonstrated why online poker security must be as technologically advanced as strategic tools are.

The main achievement of DeepStack wasn't its ability to compute poker combinations faster — computers had managed that long ago. Its true success lay in decision-making without complete information. It couldn’t see opponents’ cards, parse the entire decision tree, or have answers ready for every conceivable situation.

Instead, it continuously analyzed current play, worked with ranges, solved immediate decisions, and estimated the distant future using a neural network. It bluffed without emotion and remained unpredictable without human psychology. Poker was merely the test environment. The real question was whether AI could make quality decisions in a world where not all facts are known, and the other side harbors hidden designs. DeepStack showed the answer could be yes. And poker pros across the virtual felt were among the first to see it.

 

Sources - YouTube, Wikipedia, Science.org, onlinegamblingwebsites