Welcome to visit Communications in Theoretical Physics,
Topical Review:Statistical Physics, Soft Matter and Biophysics

A brief review of evolutionary game dynamics in the reinforcement learning paradigm

  • Guozhong Zheng 1, 2 ,
  • Xin Ou 2 ,
  • Shengfeng Deng 2 ,
  • Jiqiang Zhang , 3, * ,
  • Li Chen , 2, *
Expand
  • 1School of Physical Science and Technology, Inner Mongolia University, Hohhot 010021, China
  • 2School of Physics and Information Technology, Shaanxi Normal University, Xi'an 710061, China
  • 3School of Physics, Ningxia University, Yinchuan 750021, China

*Authors to whom any correspondence should be addressed.

Received date: 2026-01-15

  Accepted date: 2026-03-11

  Online published: 2026-04-09

Copyright

© 2026 Institute of Theoretical Physics CAS, Chinese Physical Society and IOP Publishing. All rights, including for text and data mining, AI training, and similar technologies, are reserved.
This article is available under the terms of the IOP-Standard License.

Abstract

Cooperation, fairness, trust, and resource coordination are cornerstones of modern civilization, yet their emergence remains inadequately explained, largely due to persistent discrepancies between theoretical predictions and behavioral experiments. Part of this gap may arise from the imitation learning paradigm commonly used in prior theoretical models, which assumes individuals merely copy successful neighbors according to predetermined, fixed rules. This review examines recent advances in evolutionary game dynamics that employ reinforcement learning (RL) as an alternative paradigm. In RL, individuals learn through trial and error and introspectively refine their strategies based on environmental feedback. We begin by introducing key concepts in evolutionary game theory and the two learning paradigms, then synthesize progress in applying RL to elucidate cooperation, trust, fairness, optimal resource coordination, and ecological dynamics. Collectively, these studies indicate that RL offers a promising unified framework for understanding the diverse social and ecological phenomena observed in human and natural systems.

Cite this article

Guozhong Zheng , Xin Ou , Shengfeng Deng , Jiqiang Zhang , Li Chen . A brief review of evolutionary game dynamics in the reinforcement learning paradigm[J]. Communications in Theoretical Physics, 2026 , 78(6) : 067601 . DOI: 10.1088/1572-9494/ae503e

1. Introduction

Cooperation [1], trust [2], fairness [3], and related traits are fundamental to modern society. Understanding their mechanisms is crucial for social stability, human well-being, and fostering a harmonious shared future. From a statistical physics perspective, human society can be viewed as a many-body system, where individuals act as particles and socio-economic activities form the complex interactions among them. Thus, these traits can be viewed as emergence, where we can readily borrow concepts and methods from phase transitions and critical phenomena. A notable example is resource allocation [4], which was formulated as the minority game (MG) and is elegantly solved by statistical physicists [5].
Over recent decades, evolutionary game theory (EGT) [6, 7] has advanced our understanding of how such traits emerge. By studying the evolution of prototypical games, theoretical studies have identified key mechanisms underlying these behaviors [8]. However, the rise of behavioral economics has revealed persistent inconsistencies between theoretical predictions and experimental observations [9, 10]. A significant reason for this gap may be the widespread use of the imitation learning (IL) paradigm in theoretical models [11, 12], which assumes individuals simply copy the strategies of more successful neighbors—an assumption often contradicted by experimental evidence [13]. Real-world decision-making is far more complex than the rigid logic offered by imitation.
To address these inconsistencies, researchers are increasingly adopting a fundamentally different approach: the reinforcement learning (RL) paradigm [14]. Within this paradigm, individuals introspectively optimize their strategies through interaction with environments. Crucially, RL emphasizes long-term payoff maximization, contrasting sharply with the short-term, copy-based logic of IL, where imitation may not lead to better outcomes and bypasses individual cognitive reasoning.
This brief review focuses on the RL paradigm and examines recent progress in decoding key social traits. In section 2, we introduce the framework of EGT and compare the two learning paradigms, highlighting the limitations of IL. Sections 37 present advances in applying RL to cooperation, trust, fairness, resource allocation, and ecological systems, respectively. A summary and future perspectives are provided in section 8.

2. Fundamentals

2.1. Evolutionary game theory

EGT, the core framework of this field, was introduced by John Maynard Smith and George R. Price in their seminal work, The Logic of Animal Conflict [6]. By integrating classical game theory with Darwinian evolution, EGT examines how strategies evolve and stabilize within populations, providing a versatile framework for analyzing evolutionary processes in economics, sociology, anthropology, ecology, and beyond.
A central concept in EGT is the evolutionarily stable strategy, analogous to the Nash equilibrium in classical game theory, which describes a strategy resistant to invasion by alternatives under natural selection. In well-mixed populations, the replicator equation [15] describes how strategy frequencies change over time, with each strategy's growth rate proportional to its relative payoff compared to the population average. This formulation captures strategic competition and links evolutionary dynamics directly to the game's payoff structure. EGT has since been widely applied across disciplines, including the study of altruistic behaviors discussed here.

2.2. The paradigm of imitation learning

A key component of EGT is the selection process, where better-performing strategies thrive while others diminish. Most theoretical studies model this process using IL, a paradigm in which individuals copy the strategies of more successful peers (figure 1 Left). Common implementations include the Moran process, the ‘follow-the-best' rule, and the Fermi updating rule, among others [16].
Figure 1. Two paradigms for game evolution. In imitation learning, players use the rewards of their neighbors as the utility and adopt the strategy of the neighbors who have higher utilities. Instead, players with reinforcement learning score different actions and probabilistically choose an action each time based on these scores, which are continuously revised according to the outcome.
In IL, individuals are relieved of the need to process complex environmental information, making it a simplified representation of human learning—and theoretically, an efficient mechanism for strategy updating. Empirical evidence from behavioral experiments supports this view [10], showing that approximately 62% of strategic choices can be attributed to imitative behavior. This tendency is particularly pronounced in heterogeneous environments characterized by strategic diversity, highlighting both the prevalence and significance of imitation in real-world human decision-making. Furthermore, the concept of imitation translates naturally to ecological contexts, where more adaptive individuals pass on their strategies through genetic inheritance.
Nevertheless, IL also exhibits notable limitations. Up to 38% of strategic changes remain unexplained by imitation alone [10], suggesting that human decision-making entails greater complexity than IL typically assumes. This limitation is especially evident in social systems, where individuals do not simply imitate others—particularly when interests are in conflict. Recent behavioral experiments [17] confirm that people often base decisions on others' actions rather than their payoffs, which runs counter to the core assumption of IL. This finding aligns with our intuition that we observe what others do, but not necessarily what they earn. Such discrepancies may help explain why many theoretical predictions fail to align with experimental observations [13].
In essence, IL can be regarded as a simple form of social learning [18], where individuals learn from others through observation or instruction in socio-economic activities, which may or may not involve physical practice or direct experience. This means learning occurs by observing the behaviors of others, a manner of ‘looking outward'.

2.3. The paradigm of reinforcement learning

In contrast, RL [14] offers a fundamentally different paradigm. Here, individuals learn through direct interaction with their environment, continuously refining their strategies via trial and error (figure 1 Right). RL emphasizes long-term cumulative reward maximization by balancing experience, immediate payoff, and future expectations. Rooted in psychology and neuroscience [19], RL captures how living organisms learn from experience.
The RL framework consists of four core elements: a policy (the player's decision rule), a reward (immediate performance feedback), a value function (estimating long-term returns), and the environment. Player-environment interactions typically follow the Markov property, forming a Markov decision process [20], which provides the theoretical foundation for RL.
An early RL model is the Bush–Mosteller model [21], where players follow the Pavlovian conditioned response, and those winning actions are reinforced. A more influential approach is Q-learning [22], a model-free, value-based algorithm that learns a Q-value function for state-action pairs to guide decision-making. Using an off-policy update rule, it efficiently explores the optimal policy in discrete action spaces.
In Q-learning, a Q-table stores values Q(s, a), representing the expected cumulative reward for taking action a in state s. The larger the value Q(s, a) within state s, the more preferred action a is. In the ε-greedy Q-learning, players independently choose a random action with probability ε; otherwise, they select the action with the largest Q-value within the row corresponding to their states. Afterwards, the Q-values are updated at the end of each round via the Bellman equation [14]:
$\begin{eqnarray}\begin{array}{l}Q\left({s}_{t},{a}_{t}\right)\leftarrow (1\,-\,\alpha )Q\left({s}_{t},{a}_{t}\right)\\ +\alpha \left[{{\rm{\Pi }}}_{t+1}\,+\,\gamma \mathop{\max }\limits_{{a}^{{\prime} }}Q\left({s}_{t+1},{a}^{{\prime} }\right)\right].\end{array}\end{eqnarray}$
Here, α is the learning rate, which determines how much the old experience Q(stat) is removed—the smaller the value, the more experience is retained. Πt+1 is the immediate reward obtained by taking action a in state s at time t. γ is the discount factor determining the impact of the optimal action in the next step t + 1 that one can expect. A larger γ means the player values more rewards in the future, having a long-term vision. New values integrate the experience in the past, the reward in the current step, and guidance from the future. This unambiguous interpretation makes Q-learning a mainstream choice in RL applications.
While actions are often fixed in game-theoretic contexts, the definition of the state is highly flexible and central to the functioning of RL. The power of RL lies in its ability to condition actions on environmental information, with the state serving as the representation of that environment—it determines what information the individual perceives. As expected, the design of the state space significantly influences the effectiveness of RL. On one hand, appropriate state representations that provide sufficient environmental information are essential for good performance; on the other hand, including too much information may exceed individuals' memory and learning capacities—capacities that are often assumed to be unlimited in theoretical models but are bounded in reality. In the following sections, we explore a range of state designs, from simple self-regarding setups to other-regarding configurations that incorporate neighbor states, from symmetric to asymmetric information structures, and from precise to fuzzy information representations. These varying designs not only affect individuals' learning efficiency and the resulting behaviors but also shape the model's capacity to capture real-world complexity.
Other value-based RL variants include SARSA (an on-policy alternative) and deep Q-learning, which uses neural networks to handle large and continuous state spaces. Another category of RL is policy-based methods [23, 24], where agents directly learn what actions to take without scoring them. Actor-critic algorithms integrate both value-based and policy-based methods, achieving a balance between learning efficiency and stability. Moreover, some more advanced RL techniques, such as deep RL methods [25], the multi-agent RL [26], and inverse RL [27], among others [28], have shown great potential in handling complex game-theoretic scenarios and also deserve further exploration and attention.
In short, IL and RL represent two distinct learning logics suited to different research contexts. IL focuses on leveraging existing experience through observation, corresponding to a ‘follow-the-crowd' social heuristic. In contrast, RL adopts an introspective, experience-driven approach, allowing agents to develop strategies through their own interactions with the environment—capturing a process of introspective exploration characterized as ‘learning from consequences.' The choice between these two paradigms ultimately depends on the specific research context and the questions under investigation.

3. Cooperation

Cooperation is widespread and essential in both human societies and natural systems, and the mechanisms underlying its emergence have been extensively studied [1]. In the IL paradigm [29], a key requirement is that individuals must share their strategies and payoffs during the evolutionary process—an assumption that is often neither feasible nor realistic. Because of this limitation, game-theoretic predictions under IL frequently fail to align with behavioral experiments [13], making the emergence of cooperation a persistent challenge. Recently, RL has provided a fundamentally different approach to addressing this problem, emerging as a promising paradigm for deciphering the origins of cooperation [30].

3.1. The pairwise game

The prisoner's dilemma game (PDG) is a classic pairwise model used to study cooperation. In this two-player, two-action (2 × 2) game, each player can either cooperate or defect. Mutual cooperation yields both players a reward R, whereas mutual defection results in a punishment P. If one cooperates while the other defects, the defector receives a temptation payoff T, and the cooperator receives the sucker's payoff S. Payoffs in the PDG satisfy T > R > PS and 2R > T + S. In its standard form, the payoff matrix can be expressed as
$\begin{eqnarray}\begin{array}{r}{\rm{\Pi }}=\left[\begin{array}{cc}R & S\\ T & P\end{array}\right]=\left[\begin{array}{cc}1 & -b\\ 1+b & 0\end{array}\right],\end{array}\end{eqnarray}$
where the temptation factor b quantifies the conflict between individual and collective interests. The dilemma arises because defection yields a higher individual payoff regardless of the opponent's choice, even though mutual cooperation maximizes collective welfare.
Reference [31] represents an early application of RL to explain cooperation in the PDG, focusing on the fundamental dynamics. A key finding is that cooperation emerges when players both value past experience and adopt a long-term perspective [Region I in figure 2(a)]. Cooperation also persists at a moderate level when players learn from history, even if they are short-sighted (Region III). Notably, high inequality in payoffs occurs at the boundary between Regions I and II, while rewards are nearly equal elsewhere [see figure 2(b)]. Mechanistically, the study identified that players adopt a win-stay-lose-shift strategy to sustain cooperation and converge to some coordinated optimal modes. When players ignore experience or become myopic, these modes destabilize due to unpredictable opponent behavior, and defection prevails.
Figure 2. Emergence of cooperation in the prisoner's dilemma game. (a) The phase diagram of cooperation level within the space of learning parameters (αγ) for two players playing the game, which can be divided into three regions: high cooperation (I), full defection (II), and low cooperation (III). (b) The corresponding reward difference between two players, which is visible around the boundaries between Regions I–II and I–III. b = 0.2 in equation (2) and ε = 0.01 (Adapted from [31]).
A critical aspect of RL in modeling cooperation is the design of the state—i.e. what information players perceive. Early studies often used a self-regarding setup, where the state only included the player's own previous action. This, however, provides insufficient information to capture their surroundings. Conversely, overly detailed state representations can be infeasible and may contain redundant information, since decisions often depend on only a few pieces of key cues. Reference [32] illustrates how information perception influences outcomes in the PDG: when two players operate under different information scenarios, the resulting evolutionary dynamics vary significantly, with particularly rich behaviors emerging under asymmetric perception.
Many studies have extended the two-player PDG to multi-agent settings, such as on 2D lattices, to investigate cooperation at the population level [33, 34]. Within RL, classical mechanisms from IL have been revisited, including direct reciprocity [35], indirect reciprocity [36], spatial reciprocity [37], adaptive migration [38], reputation [39], and preferential selection [40]. Regarding spatial reciprocity, Wang et al [37] demonstrated that it emerges only when players interact and learn within overlapping local neighborhoods, highlighting the importance of coupling local interaction with local strategy learning for sustaining cooperation under RL. Furthermore, introducing third parties—such as leaders [33], exporters [41], or loners [42]—has been shown to effectively promote cooperation under RL, consistent with earlier IL findings.
Beyond these well-known factors, RL-based studies of the PDG have revealed novel phenomena [4345]. Zhang et al [43] found that moderate greediness optimally promotes cooperation, a result robust across network types. In [35], global players who consider neighborhood-wide stimuli foster stronger conditional cooperation via direct reciprocity, thereby helping to sustain cooperation. Wang et al [44] showed that Lévy noise in payoffs enhances cooperation by creating a Q-value advantage for cooperative actions—an effect absent under Gaussian or no noise. In a DQN-based updating model, [45] reported that increasing the discount factor expands cooperative clusters until full cooperation is achieved, while the temptation factor b has little influence. Recently, Su et al [46] employed multi-agent RL to discover a novel strategy termed the ‘memory-two bilateral reciprocity' strategy. This strategy not only outperforms most known strategies in pairwise interactions but also dominates in evolving populations, promoting higher levels of cooperation and social welfare. This remarkable performance has been validated through both simulations and mathematical analysis, highlighting that multi-agent RL not only serves as a strategy update mechanism but also demonstrates significant potential as a strategy discovery tool. Interestingly, the overall cooperation is also enhanced—a catalytic-like effect was also discussed in [47].

3.2. The multi-player game

Beyond the pairwise game, the public goods game (PGG) serves as a paradigmatic model of multi-player cooperation, where an arbitrary number of participants are allowed in a single game. In a typical PGG, n players may each contribute an amount a ∈ [0, 1] to a common pool. The total pool is then multiplied by a synergy factor r and is divided evenly among all players. The payoff for player i is
$\begin{eqnarray}\begin{array}{rc}{{\rm{\Pi }}}_{i} & =\frac{r}{n}\displaystyle \sum _{j=1}^{n}{a}_{j}-{a}_{i}=\hat{r}\displaystyle \sum _{j=1}^{n}{a}_{j}-{a}_{i},\end{array}\end{eqnarray}$
where ai denotes player i's contribution and $\hat{r}=\frac{r}{n}$ is the reduced synergy factor. If $\hat{r}\gt 1$, the contribution ai = 1 is expected. However, for $\hat{r}\lt 1$, the dominant strategy for a rational player is to contribute nothing, i.e. a = 0. In this case, contributing while others free-ride increases group benefit at a personal cost, leading to widespread free-riding and the tragedy of the commons [48]. In many practices, a is chosen to be discrete within {0, 1}, corresponding to defection and cooperation, respectively.
Recent studies have applied RL to explore cooperation in PGGs, offering new mechanistic insights [4953]. As with PDG research, some works examine whether cooperation mechanisms known from IL also operate under RL, including reward incentives [49], voluntary participation [50], and reputation systems [51, 54]. For example, adaptive reward schemes integrated with self-regarding Q-learning can significantly raise cooperation levels [49]. Introducing third-party ‘loners' leads to stable cooperation at high synergy factors, while defector density exhibits a non-monotonic dependence on the gain factor [50]. The PGG is also studied on higher-order networks [55], where engagement is determined by Q-learning, and active players utilize social learning to act. Reputation mechanisms on hypergraphs also promote cooperation, with learning parameters being systematically examined [51]. Other factors such as neighbor influence [54] and conformity [56] are also investigated. Altogether, these studies further enrich our understanding of PGG dynamics under RL.
Most of these studies, however, adopt a self-regarding state setup, where agents base their decisions solely on their own past actions—contrary to real-world action logic, where individuals also observe their surroundings and respond to their neighbors. Reference [52] emphasizes the importance of social information in RL, comparing three models: traditional IL-based PGG, Q-learning-based PGG, and a voluntary PGG (VPGG) with Q-learning that includes a ‘loner' strategy on a lattice (figure 3). Using Fermi updating for IL as a baseline and coarse-graining surroundings by comparing cooperator/defector counts, they show that Q-learning substantially reduces the critical synergy factor $\hat{r}$ needed for cooperation, with VPGG lowering it further. Loners suppress the spread of defectors, although cooperation exhibits non-monotonic dependence on parameters. Again, cooperation is strongest when players value past experience and maintain a long-term perspective. Reference [53] further demonstrates the value of other-regarding information on spatial hypergraphs. Interestingly, two abrupt transitions in cooperation emerge as $\hat{r}$ varies, separating three regimes: no cooperation, moderate cooperation, and high cooperation. Spatial analysis reveals a chessboard-like pattern that promotes the first transition but hinders the second. Theoretical analysis of the first transition shows that far-sighted players with low exploration rates are more likely to reciprocate cooperation, thereby facilitating its emergence.
Figure 3. Schematics of three model setups. (Left) Public goods game (PGG) with a Fermi-function update rule. (Middle) PGG with Q-learning. (Right) Volunteer public goods game (VPGG) with Q-learning. The focal individuals are indicated by gray and are surrounded by neighbors that could be cooperators (red), defectors (blue), and loners (pink). Below each panel are the corresponding strategy-update components: the Fermi function (left) and the Q-tables (middle and right). (Adapted from [52].)
Notably, the IL and RL paradigms are not mutually exclusive. Human behavior often switches between different decision logics depending on context, a diversity observed in experiments [57]. Several studies capture this by allowing mixed updating rules, yielding rich evolutionary dynamics—for example, combining social learning with self-learning [58], Q-learning with Fermi rule [59, 60], Fermi rule with tit-for-tat [61], Q-learning with tit-for-tat [47], and hard with soft conditional cooperators [62], etc. These works frequently identify an optimal mixing ratio at which cooperation exceeds levels achieved under either pure rule set.

4. Trust

Trust, as a core element of human civilization [2], is often regarded as the ‘lubricant of the social system' and plays an irreplaceable role in facilitating cooperation and fostering social coordination [63]. At the individual level, trust helps establish healthy interpersonal relationships and social networks, promoting collaborative partnerships and the formation of friendships. At the society level, it enhances well-being and quality of life. From an institutional perspective, trust between the government and the public not only strengthens citizens' understanding and identification with the political system but also ensures the effective formulation and implementation of public policies, thereby contributing to social stability and development [64, 65].
Research has often employed the trust game [66] to study the evolution of trust, in which a trustor decides whether to invest part of their endowment in a trustee with trust (T) or keep it (No trust). If they keep it, the game ends. If they invest, the amount is multiplied, and the trustee then chooses either to reciprocate (R) by returning some money or to betray (B) by keeping everything. According to the assumption of Homo economicus in the classic economics—that individuals are rational and self-interested, aiming to maximize their own payoff—the trustee should always betray, and the trustor, anticipating this, should never invest. Trust, therefore, would not be expected to arise. However, behavioral experiments using the trust game [66] reveal a fundamental contradiction: trust and trustworthiness are widely observed in humans. On average, trustors invest about 50% of their endowment, and trustees return roughly 37% of the gains [67], demonstrating the pervasiveness of trust in human interactions.
To resolve this discrepancy, prior game-theoretic studies have incorporated factors such as reputation [68], population structure [69, 70], migration mechanisms [71], third-party deposit systems [72], and incentive schemes [73] into models, showing that these can trigger the emergence of trust. Notably, these works operate within the IL paradigm, where individuals replicate successful strategies to explain the spread of trust in populations [74].
Recently, studies have shifted toward the RL paradigm and shown that endogenous factors alone are sufficient to explain the emergence of trust; no exogenous factors are needed. Reference [75] adopts the Q-learning algorithm, focusing on the compromise between short-term self-interest and long-term trust benefits in the trust game. The study mainly discussed the two-player scenarios, where the two players play the role of trustor and trustee in turn. Accordingly, each player is associated with two Q-tables to guide the decision-making for the two roles, respectively. The study revealed that when individuals appreciate both their historical experience (a small learning rate α) and the returns in the future (a large discount factor γ), a high level of trust emerges naturally. As seen in figure 4, the proportion of the trust strategy peaks in the bottom-right corner, corresponding to low α (value historic experience) and high γ (long-term perspective). Q-table analysis reveals a shift in preference from short-term gain to long-term reciprocity, illustrating how trust stabilizes over time. These findings remain robust when extended to a one-dimensional lattice population.
Figure 4. Emergence of trust. Fractions of four strategies in the trust game within the parameter space (γα) ∈ (0, 1). For instance, TB denotes a player who trusts as a trustor but betrays as a trustee. High levels of trust (TR) emerge at large γ and small α (bottom-right corner), indicating that appreciating both historical experience and long-term vision promotes trust and trustworthiness. (Adapted from [75].)
More recently, [76] expanded this research to a spatial lattice by combining Q-learning with second-order social norms, exploring the synergy between reputation and learning in trust evolution. They show that Q-learning yields a richer set of steady-state strategies than the traditional Fermi rule, and even under high dilemma intensity, group wealth can improve. Reference [77] further extends the model to higher-order networks with reputation mechanisms, which significantly enhance trust and collective wealth accumulation.
Though trust theory in the RL paradigm is still in its infancy, its logic aligns with existing behavioral evidence. For example, [78] finds that future-oriented concern is essential for sustaining trust in repeated trust games: experienced subjects show much lower trust in a definite round of repeated games than in indefinitely repeated ones, where no future rewards are expected. This forward-looking motivation—beyond the immediate payoff focus of IL—is naturally captured within the RL framework.

5. Fairness

Fairness, as a cornerstone of human society, serves as a core norm for resolving conflicts of interest, such as economic inequality, climate justice, and the allocation of public resources. To investigate the mechanism of fairness, the ultimatum game (UG) [79] is often adopted as a paradigmatic model. In this game, two players divide an amount of money. One acts as the proposer, who offers p to the other, called the responder, who has an acceptance threshold q. If pq, the money is divided as proposed; otherwise, they get nothing.
According to the assumption of Homo economicus in the classic Economy, the proposer should make the smallest non-zero offer, as the responder will accept it, since something is better than nothing. Therefore, an extremely unfair outcome is predicted. However, extensive cross-cultural experiments consistently contradict this prediction, where proposers tend to offer shares around 40%–50%, and approximately 50% of responders reject unfair offers below 20% [80, 81].
To resolve this fundamental discrepancy, most previous game-theoretic works adopted the IL paradigm [82], where the strategies of better-off peers in the UG are assumed to spread more readily. Within this paradigm, researchers have revealed a bunch of factors, such as spatial structure [8385], noise [86], reputation [87, 88], role assignment [89], and empathy [90] contribute to the emergence of fairness, providing valuable insights into the mechanisms behind its evolution [91].
Recently, some studies have turned to the RL paradigm to decode the emergence of fairness [92, 93], and they have shown that the endogenous incentive is sufficient to drive the emergence of fairness and exogenous factors are not needed. In [92], they consider a two-player scenario, where the decision-making is empowered by the Q-learning algorithm. The two players take turns to play the role of proposer and responder, and therefore each is associated with two Q-tables to guide each role. The state consists of their strategy combination in the previous round, and three fixed actions for their choices for either role: low (plql < 0.5), middle (pmqm = 0.5), or high (phqh > 0.5) options. They reveal that when individuals value both historical experience (small α) and future rewards (large γ), fairness emerges significantly. An important result is that as the offer increases, the successful deals also rise, in line with observations in behavioral experiments, see figure 5. These results remain robust for different role assignments for the two players, such as rotating, random, fixed role, and for the extended scenario to a 1d lattice population.
Figure 5. Emergence of fairness. As in many practices of behavioral experiments, proposers p are offered with three options: mean (l  <  0.5), fair (m = 0.5), and overgenerous (h  >  0.5), and the responders q have the same acceptance threshold {l, m, h}. The two subplots show different dependencies of fairness on the l and h. While the rational option fractions pl and ql rise as l increases, the marginal impact of h is observed. In all cases, the densities of overgenerous options are always vanishing. Parameters: ε = 0.01, α = 0.1, γ = 0.9, h = 0.8 in (a) and l = 0.3 in (b). (Adapted from [92].)
A different implementation is given in [93], where they study the evolution of a 2d spatial UG with a strategy-adjustment Q-learning. Instead of fixed actions, the action set is composed of offer/acceptance threshold increase, decrease, or maintenance, with two sensitivity factors controlling the magnitude of adjustment. They reveal that when the two factors become imbalanced, a promoted fairness is seen. By comparison to the implementation of IL, people empowered by their Q-learning have a higher level of fairness. The study also examined the impact of learning parameters, and they reached a consistent conclusion that the appreciation of historic experience and the future rewards generally yields a high level of fairness.
The mechanism analysis for the emergence of fairness is also conducted in [92], which consists of two phases. In the first stage, those strategies leading to failed deals are removed from the system, and in the second stage, the remaining ones evolve either into the fair strategy (pmqm) or the rational strategy (plql) by a branching process. In short, the historical experience enables players to draw lessons from the past, and the expectation in future reward encourages responders to shift from a low offer to pursuing a higher fair offer. This shift may cause some immediate loss, but responders can obtain higher accumulated rewards in the long term by forcing proposers to raise their offer to reach deals.

6. Resource allocation

Resource allocation is a fundamental issue in both nature and human societies, and its efficiency directly affects system stability and sustainability [94]. Although general equilibrium theory in economics assumes that supply and demand can reach an optimal allocation, how such an optimum is achieved remains unclear. The key question is: how can populations reach an optimal allocation when individuals act in their own self-interest?
The MG [5, 95] provides perhaps the simplest toy model for studying this question, inspired by the El Farol bar problem [96]. Its core logic follows the ‘minority-wins' rule: an odd number N of individuals repeatedly choose between two options (e.g. go to the bar or stay home), and those in the minority win. In the seminal work [95], each agent is assigned a set of predefined, static strategies from a shared pool, and a phase transition is observed as model parameters vary. Many follow-up studies have explored coordination mechanisms [5, 97], and various model variants are proposed [98, 99]. A major limitation, however, is that these models rely on fixed strategies and cannot capture real-world adaptive decision-making.
Recently, RL has brought new vitality and fresh perspectives to this field [100106], primarily along two lines: value-based RL (e.g. Q-learning) and policy-based RL (e.g. REINFORCE [23]). Along the value-based line, Q-learning has shown that optimal coordination can emerge under the RL paradigm [103, 104]. [104] studies the MG in the El Farol bar context using Q-learning, where the state is the number of people who went to the bar in the last round, and actions are ‘go' or ‘not go.' Instead of ε-greedy selection, they use a softmax version with a temperature-like parameter to balance exploitation and exploration (figure 6). At low temperatures, the system gets stuck in partially coordinated local optima; at moderate temperatures, optimal coordination emerges when players value both experience and future rewards; at high temperatures, coordination breaks down into anti-coordination, and resource-utilization efficiency drops below the purely random (coin-flip) baseline. Mechanism analysis reveals there is a symmetry-breaking in action preference, where most people's preferences are stabilized, while one ‘pathetic individual' keeps switching, benefiting all others except itself.
Figure 6. Emergence of coordination in the minority game. The volatility σ2/N (a measure of coordination; lower values indicate better coordination) is plotted as a function of temperature τ, where higher τ increases the likelihood of random exploration over exploitation. Results are averaged over 100 realizations, with data collected over 5 × 103 time steps after a transient period of 5 × 104 Monte Carlo steps. The inset shows the same data on a logarithmic x-axis, with the shaded region indicating the parameter range where optimal coordination is achieved. The red dashed line at σ2/N = 0.25 corresponds to the benchmark case in which players decide randomly (e.g. by coin flip). (Adapted from [104].)
In fact, the original MG scheme [95] can be seen as a crude form of RL, where high-scoring strategies are reinforced over time. Given this fact, [103] explores the synergy between dual RL schemes—some players use classical static strategies, others use Q-learning. The study finds an optimal mixing proportion that maximizes resource allocation, marked by a first-order phase transition. The Q-learning population further self-organizes into internally and externally coordinating clusters. The latter develop momentum strategies similar to those in financial markets, which prevent long-term idling of resources but also yield lower long-term returns for those using them. This work reveals how strategy-level coordination emerges through inter-population heterogeneity.
Note that there are some early attempts [100, 101, 107] applying RL to solve the MG and claimed that the herding effect is suppressed, but with large persistent fluctuations. The reason lies in their self-regarding setup, where individuals only focus on their own actions and neglect others' choices, leaving them without enough information to coordinate effectively.
Along the second line, the policy-based methods optimize action probabilities directly without estimating a value function, and can also achieve coordination. [106] introduces a modified REINFORCE algorithm into the MG, using continuous policy probabilities and removing future reward estimation—relying only on historical payoffs. This preserves the ‘inductive learning' nature of the original MG, rather than the deductive logic of Q-learning. Their results show a coordination mechanism without symmetry breaking: individual behavior remains nearly stochastic, yet system-wide volatility stays low due to weak anticorrelation at the collective level. This symmetry-preserving coordination is further explored from a network perspective in [105].

7. Ecological systems

While RL has proven effective in decoding several human behaviors, recent studies have extended this paradigm to ecological systems. The underlying rationale is that individuals—including non-human species—actively make decisions to better adapt to their environments, which applies to most non-human species.
Much of this research focuses on predator-prey systems [108111], where Q-learning is employed to examine how individuals enhance their survival through learning. These works reveal that while predator learning tends to stabilize ecosystems, prey learning often induces oscillations in species densities—sometimes even triggering system collapse. A recent study [112] further demonstrates that survival pressure in RL-driven predator—prey systems can lead to the emergence of swarming behavior. Similarly, a flocking model based on neighbor-loss minimization [113] reproduces polarized swarming akin to phenomena observed in the Vicsek model.
Another research direction addresses biodiversity—a central theme in theoretical ecology. The rock-paper-scissors (RPS) game serves as a canonical model of cyclic dominance, yet in spatial settings, coexistence is not guaranteed. In the seminal spatial RPS model by Reichenbach et al [114], three species reproduce, predate, and migrate on a two-dimensional domain, usually called the RMF model. They coexist via spiral waves at low mobility, but go extinct when mobility exceeds a threshold—a prediction at odds with real-world observations of highly mobile species coexisting in nature.
To resolve this discrepancy, a recent spatial RPS model incorporates a joint Q-learning algorithm [115], where individuals of the same species share a common Q-table, updated collectively over generations—mirroring natural collective intelligence. Rewards are tied to survival and predation success. This framework shows that with Q-learning-guided migration, extinction becomes rare even under high mobility [figure 7(a)]. Analysis reveals that individuals develop two key tendencies: escaping predators and staying near prey. These behavioral heterogeneities suppress spiral wave formation and damp density oscillations, thereby promoting coexistence [figures 7(b), (c)]. However, the imbalance between these tendencies undermines behavioral heterogeneity and jeopardizes stable coexistence.
Figure 7. Comparison study for species coexistence with the traditional RMF and Q-learning. (a) Extinction probability versus the baseline mobility M0. The blue squares and red circles represent the results for the traditional RMF model and the Q-learning model, respectively. The extinction probability drops significantly after applying the Q-learning algorithm, greatly enhancing system stability. Typical time series of three species densities as well as the density of empty sites for the traditional RMF model (b) and our model (c), both at M0 = 3 × 10−4. This means that when species are empowered with reinforcement learning, they are much better at coexisting with each other than the RMF model, and the density oscillation is suppressed. N = 100 × 100. (Adapted from [115].)
Beyond these lines of inquiry, deep RL has also been applied to collaborative hunting [116], revealing that sophisticated coordination can emerge without high-level cognition. Other studies integrate RL with realistic ecosystems to reproduce empirical behavioral and population patterns [117], paving the way for predicting ecosystem resilience and tipping points. Additional work demonstrates RL-enabled path planning in complex settings [118], and successful navigation for microswimmers in noisy environments [119]. Together, these efforts underscore the significant potential of RL—both theoretical and experimental—in deciphering and predicting ecological dynamics.

8. Concluding remarks

Compared with mainstream IL paradigms, RL offers a fundamentally different and introspective approach to understanding behavior—one that provides a novel perspective on the origins of many human traits. Central to RL is its long-term orientation: individuals make decisions based on cumulative future rewards rather than immediate gains.
Within this framework, the above work demonstrates how four key human behaviors—cooperation, trust, fairness, and optimal resource allocation—can be explained in a unified manner. These emerge naturally when players value experience and maintain long-term objectives, without relying on external assumptions. Beyond human behaviors, the RL paradigm is also applicable to ecological systems, illustrating how it accounts for species coexistence.
Collectively, these achievements position RL as a potential unifying theoretical framework applicable to a variety of complex systems where agents—whether human or non-human—are capable of making active decisions. Moreover, RL introduces novel elements such as Q-tables, offering a unique lens through which to understand the psychological evolutionary processes underlying behavior. More critically, the value of RL extends beyond iteratively optimizing existing strategies based on cumulative rewards; its capacity to autonomously explore and discover optimal strategies that surpass human-prescribed designs further expands the boundaries of this paradigm [46]. This endows RL with a dual function—both updating existing strategies and discovering new ones—thereby providing richer research perspectives and methodological support for subsequent analyses of behavioral evolution in various complex systems. It is important to note that RL is not a superior replacement of IL; instead, the two paradigms are complementary and are a context-dependent choice for the system under study.
However, this promising picture should not obscure the fact that several fundamental challenges remain. As research progresses, a series of key questions warranting further exploration has come to the fore, particularly regarding the realism of state representations and the experimental validation of RL paradigms in real-world human decision-making.
First, realistically portraying state representation constitutes a core challenge—how to more authentically reconstruct the information environment in which individuals make decisions within models remains an inadequately addressed issue. Most current research pays insufficient attention to the cognitive foundations and informational constraints underlying state settings. In real-world scenarios, individuals often face objective conditions such as incomplete information, limited perceptual capabilities, and cognitive load constraints. Furthermore, significant heterogeneity exists across individuals in terms of information access channels and processing methods. State design is thus perpetually confronted with a dilemma: incorporating excessive information can easily lead to dimensional explosion, substantially increasing learning complexity; whereas overly simplified information may restrict an individual's ability to explore and acquire optimal strategies. Consequently, investigating how to achieve effective dimensionality reduction of information and how to systematically integrate more realistic cognitive constraints into the modeling process will provide crucial support for narrowing the gap between theoretical models and real-world decision-making scenarios.
Second, although RL has demonstrated immense potential as a unified theoretical paradigm, its core foundational principles still lack sufficient direct validation through behavioral experiments—a fact that significantly impedes the development of a robust theoretical framework. It is worth noting that existing relevant behavioral experimental research [120123] has already accumulated rich experimental paradigms and empirical evidence. These data can serve as resources for subsequent validation of the foundational principles of RL and optimization of state representations.
Third, while RL has demonstrated success in replicating experimental observations, it is important to note that these explanatory capabilities are not unique to this paradigm; other bounded rationality models, and adaptive learning models may likewise reproduce similar behavioral patterns. Therefore, a systematic comparison between RL and alternative learning rules—evaluated from the dual perspectives of experimental fit and theoretical predictive power—represents an important and promising direction for future research.

We would like to thank Professor Weiran Cai and all other collaborators who have contributed to the research on this topic. This work is supported by the National Natural Science Foundation of China (Grants Nos. 12075144, 12165014), the Fundamental Research Funds for the Central Universities (Grant No. GK202401002), and the Key Research and Development Program of Ningxia Province in China (Grant No. 2021BEB04032).

1
Nowak M A >2006 Five rules for the evolution of cooperation Science 314 1560

DOI

2
Hardin R >2002 Trust and Trustworthiness Russell Sage Foundation

3
Piketty T >2014 Capital in the Twenty-First Century Belknap Press: An Imprint of Harvard University Press

4
Arthur W B >1999 Complexity and the Economy Science 284 107

DOI

5
Challet D, Marsili M, Zhang Y C >2005 Minority Games: Interacting Agents in Financial Markets OUP Oxford

6
Smith J M, Price G R >1973 The logic of animal conflict Nature 246 15

DOI

7
Smith J M >1982 Evolution and the Theory of Games Cambridge University Press

8
Perc M, Jordan J J, Rand D G, Wang Z, Boccaletti S, Szolnoki A >2017 Statistical physics of human cooperation Phys. Rep. 687 1

DOI

9
Camerer C F >2011 Behavioral Game Theory: Experiments in Strategic Interaction Princeton University Press

10
Traulsen A, Semmann D, Sommerfeld R D, Krambeck H-J, Milinski M >2010 Human strategy updating in evolutionary games Proc. Natl. Acad. Sci. USA 107 2962

DOI

11
Nowak M A, May R M >1992 Evolutionary games and spatial chaos Nature 359 826

DOI

12
Szabó G, Töke C >1998 Evolutionary prisoner's dilemma game on a square lattice Phys. Rev. E 58 69

DOI

13
Sánchez A >2018 Physics of human cooperation: experimental evidence and theoretical models J. Stat. Mech.: Theory Exp. 2018 024001

DOI

14
Sutton R S, Barto A G >2018 Reinforcement Learning: An Introduction MIT Press

15
Taylor P D, Jonker L B >1978 Evolutionary stable strategies and game dynamics Math. Biosci. 40 145

DOI

16
Szabó G, Fáth G >2007 Evolutionary games on graphs Phys. Rep. 446 97

DOI

17
Grujić J, Gracia-Lázaro C, Milinski M, Semmann D, Traulsen A, Cuesta J A, Moreno Y, Sánchez A >2014 A comparative analysis of spatial Prisoner's Dilemma experiments: conditional cooperation and payoff irrelevance Sci. Rep. 4 4615

DOI

18
Bandura A >1977 Social Learning Theory Prentice-hall Englewood Cliffs, NJ

19
Lee D, Seo H, Jung M W >2012 Neural basis of reinforcement learning and decision making Annu. Rev. Neurosci. 35 287

DOI

20
Puterman M L >2014 Markov Decision Processes: Discrete Stochastic Dynamic Programming Wiley

21
Bush R R, Mosteller F >1955 Stochastic Models for Learning Wiley

22
Watkins C J C H >1989 Learning from delayed rewards PhD Thesis University of Cambridge

23
Williams R J >1992 Simple statistical gradient-following algorithms for connectionist reinforcement learning Mach. Learn. 8 229

DOI

24
Sutton R S, McAllester D, Singh S, Mansour Y >1999 Advances in Neural Information Processing Systems MIT Press

25
Vincent F-L, Henderson P, Islam R, Bellemare M G, Pineau J >2018 Found. Trends Mach. Learn. 11 219 354

DOI

26
Albrecht S V, Christianos F, Schäfer L >2024 Multi-agent Reinforcement Learning: Foundations and Modern Approaches MIT Press

27
Arora S, Doshi P >2021 A survey of inverse reinforcement learning: challenges, methods and progress Artif. Intell. 297 103500

DOI

28
Wang Z, Mu C, Hu S, Chu C, Li X >2022 Modelling the dynamics of regret minimization in large agent populations: a master equation approach Proc. IJCAI 534 540

29
Smith J M >1982 Did Darwin get it Right? Essays on games, sex and evolution Springer p 202

30
Xie K, Szolnoki A >2026 Reinforcement learning in evolutionary game theory: a brief review of recent developments Appl. Math. Comput. 510 129685

DOI

31
Ding Z, Zheng G, Cai C, Cai W, Chen L, Zhang J, Wang X >2023 Emergence of cooperation in two-agent repeated games with reinforcement learning Chaos, Solitons Fractals 175 114032

DOI

32
Zheng G, Ding Z, Zhang J, Deng S, Cai W, Chen L >2025 Evolution of cooperation with Q-learning: the impact of information perception Chaos 35 053129

DOI

33
Ding H, Zhang G, Wang S, Li J, Wang Z >2019 Q-learning boosts the evolution of cooperation in structured population by involving extortion Physica A 536 122551

DOI

34
Lee H, Chen S, Shi F >2025 Enhancing cooperation in dynamic networks through reinforcement-learning-based rewiring strategies New J. Phys. 27 013025

DOI

35
Jia D, Guo H, Song Z, Shi L, Deng X, Perc M, Wang Z >2021 Local and global stimuli in reinforcement learning New J. Phys. 23 083020

DOI

36
Zhao C, Zheng G, Zhang C, Zhang J, Chen L >2024 Emergence of cooperation under punishment: a reinforcement learning perspective Chaos 34 073123

DOI

37
Wang L, Shi X, Zhou Y >2025 Spatial reciprocity under reinforcement learning mechanism Chaos 35 023103

DOI

38
Fang Z, Xu H, Xie C, Yue X, Benko T P, Huang C >2025 Evolution of cooperation in multi-agent systems driven by reputation-based migration Chaos, Solitons Fractals 200 117115

DOI

39
Zhang Q, Yan Y >2025 Cooperation enhancement through a double-layer coupling mechanism with varying interaction radius in Prisoner's Dilemma Phys. Lett. A 2025 130754

DOI

40
Bai P, Qiang B, Zou K, Huang C >2024 Preferential selection based on adaptive attractiveness induce by reinforcement learning promotes cooperation Chaos, Solitons Fractals 180 114592

DOI

41
You T, Yang H, Wang J, Zhang P, Chen J, Zhang Y >2023 Cooperative behavior under the influence of multiple experienced guiders in prisoner's dilemma game Appl. Math. Comput. 458 128234

42
Huang Y, Chen Y >2025 Promoting cooperation in the voluntary prisoner's dilemma game via reinforcement learning Chaos 35 043130

DOI

43
Zhang H-F, Wu Z-X, Wang B-H >2012 Universal effect of dynamical reinforcement learning mechanism in spatial evolutionary games J. Stat. Mech.: Theory Exp. 2012 P06005

DOI

44
Wang L, Jia D, Zhang L, Zhu P, Perc M, Shi L, Wang Z >2022 Lévy noise promotes cooperation in the prisoner's dilemma game with reinforcement learning Nonlinear Dyn. 108 1837

DOI

45
Wang X, Yang Z, Liu Y, Chen G >2023 A reinforcement learning-based strategy updating model for the cooperative evolution Physica A 618 128699

DOI

46
Su Q, Wang H, Xia Y, Wang L >2025 A multi-agent reinforcement learning framework for exploring dominant strategies in iterated and evolutionary games Nat. Commun. 17 490

DOI

47
Sheng A, Zhang J, Zheng G, Zhang J, Cai W, Chen L >2024 Catalytic evolution of cooperation in a population with behavioral bimodality Chaos 34 103117

DOI

48
Hardin G >1968 The tragedy of the commons Science 162 1243

DOI

49
Wang L, Fan L, Zhang L, Zou R, Wang Z >2023 Synergistic effects of adaptive reward and reinforcement learning rules on cooperation New J. Phys. 25 073008

DOI

50
Zhang H, An T, Yan P, Hu K, An J, Shi L, Zhao J, Wang J >2024 Exploring cooperative evolution with tunable payoff's loners using reinforcement learning Chaos, Solitons Fractals 178 114358

DOI

51
Zou K, Huang C >2024 Incorporating reputation into reinforcement learning can promote cooperation on hypergraphs Chaos, Solitons Fractals 186 115203

DOI

52
Zheng G, Zhang J, Deng S, Cai W, Chen L >2024 Evolution of cooperation in the public goods game with Q-learning Chaos, Solitons Fractals 188 115568

DOI

53
Li B, Zhang Z, Zheng G, Cai C, Zhang J, Chen L >2025 Cooperation in public goods games: leveraging other-regarding reinforcement learning on hypergraphs Phys. Rev. E 111 014304

DOI

54
Kang H, Jiang C, Shen Y, Sun X, Chen Q >2025 Neighbor-aware reinforcement learning fosters cooperation in spatial public goods games Chaos, Solitons Fractals 199 116862

DOI

55
Xu Y, Wang J, Chen J, Zhao D, Özer M, Xia C, Perc M >2024 Reinforcement learning and collective cooperation on higher-order networks Knowl.-Based Syst. 301 112326

DOI

56
Zhang L, Li Y, Xie Y, Feng Y, Huang C >2025 The combined effects of conformity and reinforcement learning on the evolution of cooperation in public goods games Chaos, Solitons Fractals 193 116071

DOI

57
Traulsen A, Semmann D, Sommerfeld R D, Krambeck H-J, Milinski M >2010 Human strategy updating in evolutionary games Proc. Natl. Acad. Sci. USA 107 2962

DOI

58
Han X, Zhao X, Xia H >2022 Hybrid learning promotes cooperation in the spatial prisoner's dilemma game Chaos, Solitons Fractals 164 112684

DOI

59
Zhang Y, Zheng Z, Zhang X, Ma J >2025 A layered strategy updating mechanism for spatial public goods game with punishment Chaos, Solitons Fractals 201 117264

DOI

60
Yang Y, Zhao D, Wang J >2025 Evolution of cooperation in spatial public goods games driven by reinforcement learning and environmental feedback Chaos, Solitons Fractals 199 116592

DOI

61
Ma L, Zhang J, Zheng G, Liang R, Chen L >2023 Emergence of cooperation in a population with bimodal response behaviors Chaos, Solitons Fractals 171 113452

DOI

62
Zhao C, Feng X, Zheng G, Cai W, Zhang J, Chen L >2025 Evolution of cooperation in a dual-mode mixture of conditional cooperators Phys. Rev. E 112 054309

DOI

63
Arrow K J >1974 The Limits of Organization Norton & Company

64
Zak P J, Knack S >2001 Trust and growth Econ. J. 111 295

DOI

65
Algan Y, Cahuc P >2013 Trust and growth Annu. Rev. Econ. 5 521

DOI

66
Berg J, Dickhaut J, McCabe K >1995 Trust, reciprocity, and social history Games Econ. Behav. 10 122

DOI

67
Johnson N D, Mislin A A >2011 Trust games: a meta-analysis J. Econ. Psychol. 32 865

DOI

68
Bravo G, Tamburino L >2008 The evolution of trust in non-simultaneous exchange situations Rationality Soc. 20 85

DOI

69
Wang C >2024 Evolution of trust in structured populations Appl. Math. Comput. 471 128595

DOI

70
Guo R, Liu L, Liu Y, Zhang L >2023 Evolution of trust in a hierarchical population with different investors based on investment behavioral theory Chaos, Solitons Fractals 176 114078

DOI

71
Zhu Y, Li W, Xia C, Chica M >2024 Payoff-driven migration promotes the evolution of trust in networked populations Knowl.-Based Syst. 305 112645

DOI

72
Guo R, Liu L, Liu Y, Zhang L >2024 Evolution of trust in the N-player trust game with the margin system Appl. Math. Comput. 473 128649

DOI

73
Liu Y, Wang L, Guo R, Hua S, Liu L, Zhang L, Han T A >2025 Evolution of trust in the N-player trust game with transformation incentive mechanism J. R. Soc. Interface 22 20240726

DOI

74
Kumar A, Capraro V, Perc M >2020 The evolution of trust and trustworthiness J. R. Soc. Interface 17 20200491

DOI

75
Zheng G, Zhang J, Zhang J, Cai W, Chen L >2024 Decoding trust: a reinforcement learning perspective New J. Phys. 26 053041

DOI

76
Zhu Y, Xing B, Xia C >2025 Q-learning update with second-order reputation promotes the evolution of trust within structured populations Chaos, Solitons Fractals 199 116653

DOI

77
Hu Z, Zhu Y, Zhao D, Xia C >2026 The higher-order networked N-player trust game driven by reputation and reinforcement learning Chaos, Solitons Fractals 202 117623

DOI

78
Engle-Warnick J, Slonim R L >2004 The evolution of strategies in a repeated trust game J. Econ. Behav. Organ. 55 553

DOI

79
Güth W, Schmittberger R, Schwarze B >1982 An experimental analysis of ultimatum bargaining J. Econ. Behav. Organ. 3 367

DOI

80
Thaler R H >1988 Anomalies: the ultimatum game J. Econ. Perspect. 2 195

DOI

81
Güth W, Kocher M G >2014 More than thirty years of ultimatum bargaining experiments: motives, variations, and a survey of the recent literature J. Econ. Behav. Organ. 108 396

DOI

82
Szabó G, UQke C >1998 Evolutionary prisoner's dilemma game on a square lattice Phys. Rev. E 58 69

DOI

83
Page K M, Sigmund K >2000 The spatial ultimatum game Proc. Biol. Sci. 267 2177

DOI

84
Kuperman M N, Risau-Gusman S >2008 The effect of the topology on the spatial ultimatum game Eur. Phys. J. B 62 233

DOI

85
Iranzo J, Román J M, Sánchez Á >2011 The spatial ultimatum game revisited J. Theor. Biol. 278 1

DOI

86
Gale J, Binmore K G, Samuelson L >1995 Learning to be imperfect: the ultimatum game Games Econ. Behav. 8 56

DOI

87
Zhang Y, Yang S, Chen X, Bai Y, Xie G >2023 Reputation update of responders efficiently promotes the evolution of fairness in the ultimatum game Chaos, Solitons Fractals 169 113218

DOI

88
Deng L, Li W, Wang R, Wang C >2025 The impact of reputation-based dynamic reward mechanism on the evolution of fairness Chaos, Solitons Fractals 199 116861

DOI

89
Yang Z >2023 Role polarization and its effects in the spatial ultimatum game Phys. Rev. E 108 024106

DOI

90
Page K M, Nowak M A >2002 Empathy leads to fairness Bull. Math. Biol. 64 1101

DOI

91
Debove S, Baumard N, André J-B >2016 Models of the evolution of fairness in the ultimatum game: a review and classification Evol. Hum. Behav. 37 245

DOI

92
Zheng G, Zhang J, Ou X, Deng S, Chen L >2025 Decoding fairness: a reinforcement learning perspective Phys. Rev. E 111 064307

DOI

93
Wu B, Shen S, Wang J, Wan H >2025 Q-learning promotes the evolution of fairness and generosity in the ultimatum game Chaos, Solitons Fractals 200 116984

DOI

94
Samuelson P, Nordhaus W >2005 Economics 18th edn McGraw-Hill Education

95
Challet D, Zhang Y-C >1997 Emergence of cooperation and organization in an evolutionary game Physica A 246 407

DOI

96
Arthur W B >1994 Inductive reasoning and bounded rationality Am. Econ. Rev. 84 406

97
Chakraborti A, Challet D, Chatterjee A, Marsili M, Zhang Y-C, Chakrabarti B K >2015 Statistical mechanics of competitive resource allocation using agent-based models Phys. Rep. 552 1

DOI

98
Zhou T, Wang B, Zhou P, Yang C, Liu J >2005 Self-organized Boolean game on networks Phys. Rev. E 72 046139

DOI

99
Zhang J, Huang Z, Dong J, Huang L, Lai Y-C >2013 Controlling collective dynamics in complex minority-game resource-allocation systems Phys. Rev. E 87 052808

DOI

100
Zhang S, Dong J, Zhang H, Lu Y, Wang J, Huang Z >2024 Self organizing optimization and phase transition in reinforcement learning minority game system Front. Phys. 19 40201

DOI

101
Zhang S, Dong J, Liu L, Huang Z, Huang L, Lai Y-C >2019 Reinforcement learning meets minority game: toward optimal resource allocation Phys. Rev. E 99 032302

DOI

102
Zhang S, Zhang J, Huang Z, Guo B, Wu Z, Wang J >2019 Collective behavior of artificial intelligence population: transition from optimization to game Nonlinear Dyn. 95 1627

DOI

103
Zhang Z, Zheng G, Chen L, Cai C, Deng S, Li B, Zhang J >2026 Dual reinforcement learning synergy in resource allocation: emergence of momentum strategy Chaos, Solitons Fractals 202 117441

DOI

104
Zheng G, Cai W, Qi G, Zhang J, Chen L >2025 Optimal coordination of resource: a solution from reinforcement learning Phys. Rev. E 112 064305

DOI

105
Shao C, Rao W, Xu W, Wei L >2025 Network analysis on the symmetric coordination in a reinforcement-learning-based minority game Entropy 27 676

DOI

106
Rao W, Han M, Xu W >2025 Emergent coordination without symmetry breaking in minority game via policy-based reinforcement learning Chaos, Solitons Fractals 198 116550

DOI

107
Andrecut M, Ali M K >2001 Q learning in the minority game Phys. Rev. E 64 067103

DOI

108
Olsen M M, Fraczkowski R >2015 Co-evolution in predator prey through reinforcement learning J. Comput. Sci. 9 118

DOI

109
Wang X, Cheng J, Wang L >2019 Deep-reinforcement learning-based co-evolution in a predator-prey system Entropy 21 773

DOI

110
Wang X, Cheng J, Wang L >2020 A reinforcement learning-based predator-prey model Ecol. Complex. 42 100815

DOI

111
Park J, Lee J, Kim T, Ahn I, Park J >2021 Co-evolution of predator-prey ecosystems by reinforcement learning agents Entropy 23 461

DOI

112
Li J, Li L, Zhao S >2023 Predator-prey survival pressure is sufficient to evolve swarming behaviors New J. Phys. 25 092001

DOI

113
Durve M, Peruani F, Celani A >2020 Learning to flock through reinforcement Phys. Rev. E 102 012601

DOI

114
Reichenbach T, Mobilia M, Frey E >2007 Mobility promotes and jeopardizes biodiversity in rock-paper-scissors games Nature 448 1046

DOI

115
Jiang K, Zhao C, Deng S, Cai W, Zhang J, Chen L >2025 Species coexistence in the reinforcement learning paradigm arXiv:2508.17599

116
Tsutsui K, Tanaka R, Takeda K, Fujii K >2024 Collaborative hunting in artificial agents with deep reinforcement learning eLife 13 e85694

DOI

117
Strannegård C, Palak M, Engsner N, Stocco A, Antonelli A, Silvestro D >2025 Predicting ecosystem resilience using multi-agent reinforcement learning bioRxiv 2025.06.07.658424

118
Nasiri M, Liebchen B >2022 Reinforcement learning of optimal active particle navigation New J. Phys. 24 073042

DOI

119
Muinos-Landin S, Fischer A, Holubec V, Cichos F >2021 Reinforcement learning with artificial microswimmers Sci. Robot. 6 eabd9285

DOI

120
Wang Z, Jusup M, Wang R-W, Shi L, Iwasa Y, Moreno Y, Kurths J >2017 Onymity promotes cooperation in social dilemma experiments Sci. Adv. 3 e1601444

DOI

121
Wang Z, Jusup M, Shi L, Lee J-H, Iwasa Y, Boccaletti S >2018 Exploiting a cognitive bias promotes cooperation in social dilemma experiments Nat. Commun. 9 2954

DOI

122
Wang Z >2020 Communicating sentiment and outlook reverses inaction against collective risks Proc. Natl. Acad. Sci. U.S.A. 117 17650

DOI

123
Jia D, Romic I, Shi L, Su Q, Liu C, Liu J, Holme P, Li X, Wang Z >2025 Social networking agency and prosociality are inextricably linked in economic games Nat. Hum. Behav. 9 2620

DOI

Outlines

/