| Outcome | First rounds (%) | All rounds (%) |
|---|---|---|
| SIM cooperation | 17.05 | 16.79 |
| SEQ first-mover cooperation | 30.90 | 29.63 |
| SEQ second mover after C | 62.31 | 67.21 |
| SEQ second mover after D | 3.32 | 3.37 |
| CHAT cooperation | 79.08 | 74.59 |
| SIM mutual cooperation | 4.41 | 4.99 |
| SEQ mutual cooperation | 19.25 | 19.92 |
| CHAT mutual cooperation | 69.86 | 66.37 |
In-Silico Experiments of Varieties of Repeated Games
Endogenous strategic risk, communication, and cooperation among frontier language models
Abstract. Kartal and Müller (2026) explain cooperation in indefinitely repeated dilemmas by making strategic risk endogenous: players privately differ in their utility from cooperation, so the danger that an opponent defects is itself an equilibrium object. Their model predicts that sequential moves reduce strategic risk and that honest cheap talk can reduce it further. We implement a computational analogue with five frontier language-model families—GPT, Claude, Gemini, DeepSeek, and Grok—in simultaneous, sequential, and pre-play-chat repeated prisoner’s dilemmas at horizons 10, 20, and 50. The resulting balanced round robin contains 225 ordered matches and 6,000 rounds. The qualitative institutional ranking matches the human experiment: mutual cooperation is 26.2% in simultaneous play, 36.3% in sequential play, and 85.8% with chat. Heterogeneity is stark: Claude and Grok almost entirely defect without communication, while both become highly cooperative after chat; GPT remains the most robust cooperator across institutions. This is an exploratory in-silico pilot rather than a model-comparison benchmark: there is one match per design cell, horizons are fixed rather than geometrically terminated, and provider-specific inference settings differ across families.
1. Motivation
The folk theorem gives repeated games an embarrassment of equilibria. Once players are sufficiently patient, both persistent cooperation and persistent defection can be supported, and small institutional changes often leave the equilibrium set essentially unchanged. Laboratory behavior is much sharper. Cooperation reacts to the payoff matrix, the continuation probability, the timing of moves, the composition of the subject pool, and the opportunity to communicate.
Kartal and Müller (2026) propose a disciplined way to recover those comparative statics. Instead of treating strategic risk as an external equilibrium-selection parameter, they derive it from private heterogeneity in preferences for cooperation. Their framework yields cutoff equilibria and ordered predictions across simultaneous play, sequential play, and communication. Their laboratory experiment strongly confirms the ordering.
Large language models provide a new population on which to examine the same institutions. They are not replicas of human subjects: their “preferences” are induced by training, prompting, and inference configuration rather than stable utility parameters. But they are heterogeneous, strategically responsive, and capable of natural-language communication. This makes them useful as computational subjects for asking three narrower questions:
- Do the institutional rankings from the human experiment survive with LLM agents?
- Does a longer announced horizon act like a stronger shadow of the future?
- Is cooperation a generic model capability, or does it depend on matching a model with its own family versus another family?
2. The benchmark model
2.1 Private preferences and endogenous strategic risk
The monetary stage game is
| Other cooperates | Other defects | |
|---|---|---|
| Cooperate | \(c,c\) | \(a,b\) |
| Defect | \(b,a\) | \(d,d\) |
with \(b>c>d>a\) and \(2c>a+b\). Player \(i\) has a privately observed type \(\gamma_i\), drawn independently from a commonly known distribution \(F\). If the player chooses \(C\), stage utility is monetary payoff plus \(\gamma_i\); if the player chooses \(D\), stage utility is monetary payoff alone. Thus a larger \(\gamma_i\) means a stronger preference for cooperation.
The analysis focuses on two repeated-game strategies:
- CS: cooperate while cooperation survives, then punish forever after a breakdown;
- AD: always defect.
If an opponent uses CS with probability \(p\), the net expected utility from CS rather than AD is
\[ \Pi(\delta,p,\gamma) =p\left(\frac{c+\gamma}{1-\delta}-b-\frac{\delta d}{1-\delta}\right) +(1-p)(a+\gamma-d). \]
The second term is the loss from attempting cooperation against an always-defect opponent. In equilibrium, \(p\) is not imposed by an analyst: it equals the mass of types whose own incentives lead them to choose CS. The most cooperative perfect Bayesian equilibrium has a cutoff \(\gamma^*\) satisfying
\[ \Pi\!\left(\delta,1-F(\gamma^*),\gamma^*\right)=0. \]
Types above \(\gamma^*\) choose CS and types below it choose AD. Strategic risk is therefore \(F(\gamma^*)\), the equilibrium probability of encountering an AD type.
2.2 Comparative statics
The cutoff representation generates the paper’s basic predictions. Cooperation rises when the reward from mutual cooperation \(c\), the sucker payoff \(a\), or patience \(\delta\) rises; it rises when the temptation payoff \(b\) or punishment payoff \(d\) falls; and it rises when the type distribution shifts toward stronger preferences for cooperation. The result is sharper than the classical repeated-game condition because it predicts changes in the prevalence of cooperation, not merely whether a cooperative equilibrium exists.
Sequential play has two cutoffs. A second mover who observes first-mover cooperation faces no uncertainty about the current action, so \(\gamma_S^*\) solves
\[ \frac{c+\gamma_S^*}{1-\delta}-b-\frac{\delta d}{1-\delta}=0. \]
The first mover anticipates this conditional response and therefore faces less strategic risk than under simultaneous play. Kartal and Müller show that both the first-mover cutoff \(\gamma_F^*\) and the second-mover cutoff \(\gamma_S^*\) lie below the simultaneous cutoff \(\gamma^*\). Consequently, first-mover cooperation, second-mover cooperation after an observed \(C\), continuation cooperation, and mutual cooperation are all predicted to be higher in SEQ than in SIM. The model does not order \(\gamma_F^*\) and \(\gamma_S^*\): first movers cannot earn the temptation payoff, while second movers face no strategic uncertainty after observing \(C\).
Communication adds a second private dimension: some players are honest and will not announce CS if they intend to play AD. Mutual cooperative messages then screen out at least some defectors. Any positive honesty rate raises continuation and mutual cooperation above SIM; if honesty is sufficiently prevalent, CHAT also exceeds SEQ. Under perfect honesty, nobody is cheated following a mutual cooperative agreement, and the cooperative cutoff collapses to the comparison between the streams \(c+\gamma\) and \(d\).
These results imply a clean institutional ranking:
\[ \text{mutual cooperation: CHAT} > \text{SEQ} > \text{SIM}, \]
provided communication is sufficiently honest.
3. Human evidence in Kartal and Müller
3.1 Design
The paper implements the monetary payoff matrix \((a,b,c,d)=(12,50,32,25)\) with continuation probability \(\delta=0.5\):
| Other cooperates | Other defects | |
|---|---|---|
| Cooperate | 32, 32 | 12, 50 |
| Defect | 50, 12 | 25, 25 |
The design is between subjects. SIM uses simultaneous actions; SEQ fixes subjects as first or second movers for the entire session; CHAT permits free-form communication before every simultaneous action. Each treatment contains six independent matching groups. Matches end geometrically with probability \(1-\delta=0.5\) after each round. The experiment was run at the Vienna Center for Experimental Economics in 2017 with 280 participants; subjects completed repeated matches for up to 75 minutes or 60 matches and earned €19.43 on average.
3.2 Results
The human data support both central hypotheses. First-mover cooperation in SEQ exceeds SIM, and second movers condition sharply on the observed first action: they cooperate 67.21% of the time after \(C\) but only 3.37% after \(D\). Mutual cooperation is 4.99% in SIM, 19.92% in SEQ, and 66.37% in CHAT; all pairwise differences are reported as significant at \(p<0.001\). The ordering remains when the authors restrict attention to experienced subjects.
Communication is behaviorally meaningful rather than decorative. External coders classify messages as cooperative, neutral, or defective. Mutual cooperative signals strongly predict mutual cooperation, and more than 91% of subjects follow through after exchanging cooperative messages. The paper interprets this honesty as the mechanism that lets chat screen strategic types and coordinate play.
4. In-silico design
We retain the paper’s payoff matrix but replace the human subject pool with five model families accessed through OpenRouter:
- GPT-5.6 Sol;
- Claude Sonnet 5;
- Gemini 3.6 Flash;
- DeepSeek V4 Pro;
- Grok 4.5.
Every ordered pair—including self-play—is run in three institutions and at fixed horizons \(T\in\{10,20,50\}\). This yields \(5^2\times3\times3=225\) matches and 6,000 observed rounds. The ordering is intentional: in SEQ, reversing a pair changes which family moves first; even in SIM and CHAT, keeping ordered seats exposes prompt- or role-label asymmetries.
The treatments are direct computational analogues:
- SIM: both models choose without observing the other’s current action;
- SEQ: player 1 chooses first and player 2 observes that action;
- CHAT: models exchange messages of at most 25 words, observe both messages, and then choose simultaneously.
Actions are restricted to Cooperate and Defect. The runner logs prompts, raw and parsed actions, messages, payoffs, model identities, roles, horizon, and seed after every round. Invalid outputs are retried and affected match suffixes are regenerated rather than retained. The final panel has no invalid action records.
The experiment differs from Kartal and Müller in three important ways. First, horizons are finite and announced, not geometrically terminated. Horizon length is therefore only a rough proxy for the shadow of the future; it is not a change in \(\delta\). Second, this baseline sets the private system-prompt parameter \(\gamma\) to zero for both models, so heterogeneity comes from model families rather than experimentally assigned preference types. Third, endpoint constraints required provider-specific inference settings. GPT and Claude ran without hidden reasoning; Gemini required a small reasoning allowance; DeepSeek required a larger low-effort allowance to avoid consuming the entire response budget before emitting an action; and Grok required reasoning to be enabled but was capped at 64 reasoning tokens. These choices preserve actions but weaken any interpretation of model-family coefficients as pure architectural effects.
There is one match per ordered-pair × treatment × horizon cell. Rounds within a match are serially dependent, and no independent seed replication is available. All comparisons below are descriptive pilot results, not standard errors or hypothesis tests.
5. Aggregate results
5.1 The institutional ordering survives
| Treatment | Rounds | Player 1 coop. | Player 2 coop. | Mutual coop. |
|---|---|---|---|---|
| SIM | 2000 | 28.4% | 27.9% | 26.2% |
| SEQ | 2000 | 38.0% | 36.4% | 36.3% |
| CHAT | 2000 | 86.7% | 87.5% | 85.8% |
Mutual cooperation is 26.2% in SIM, 36.3% in SEQ, and 85.8% in CHAT. Thus the LLM population reproduces the qualitative ranking \(\text{CHAT}>\text{SEQ}>\text{SIM}\). Relative to the human experiment, LLM mutual cooperation is higher by 21.2 percentage points in SIM, 16.4 points in SEQ, and 19.4 points in CHAT.
The sequential mechanism is particularly stark. Across all SEQ rounds, second movers cooperate after a first-mover \(C\) in 95.5% of observations and after a first-mover \(D\) in only 0.08%. In first rounds alone, the corresponding rates are 55.9% and 0%. Subsequent LLM play therefore locks into conditional cooperation more strongly than the first round does.
The first-mover prediction is more nuanced. Across all rounds, player-1 cooperation rises from 28.4% in SIM to 38.0% in SEQ, consistent with reduced strategic risk. In first rounds, however, player-1 cooperation is lower in SEQ (45.3%) than in SIM (57.3%). The human first-round ranking therefore does not replicate. The divergence is a warning against mapping an announced finite horizon directly into the paper’s stationary infinite-horizon cutoff model.
5.2 Horizon patterns
SIM mutual cooperation rises from 16.8% at \(T=10\) to 29.8% at \(T=50\). CHAT rises more sharply, from 67.6% to 92.2%. SEQ is non-monotone—26.0%, 21.4%, and 44.3%—but is highest at the longest horizon. The horizon gradient is therefore directionally consistent with a stronger shadow of the future in SIM and CHAT, but the finite design does not identify the comparative static in \(\delta\).
6. Head-to-head model comparisons
6.1 Self-play versus cross-family play
| Treatment | Pair type | Rounds | Mutual cooperation |
|---|---|---|---|
| CHAT | Cross-family | 1600 | 85.3% |
| CHAT | Self-play | 400 | 87.5% |
| SEQ | Cross-family | 1600 | 34.2% |
| SEQ | Self-play | 400 | 44.5% |
| SIM | Cross-family | 1600 | 23.7% |
| SIM | Self-play | 400 | 36.0% |
Self-play is more cooperative than cross-family play in every institution, but the gap depends strongly on whether models can communicate. It is 12.3 percentage points in SIM (36.0% versus 23.7%), 10.3 points in SEQ (44.5% versus 34.3%), and only 2.2 points in CHAT (87.5% versus 85.3%). Communication nearly eliminates the aggregate “same-family premium.”
That aggregate comparison is not a generic preference for one’s own family. It is driven heavily by the identity of the match. Claude’s no-chat behavior, in particular, creates zeros in every pair containing Claude. Pair-level matrices are therefore more informative than a single self-play coefficient.
6.2 Pairwise cooperation matrices
Five patterns stand out.
First, GPT is the most robust cooperator. Pooling seats and opponents, GPT cooperates in 62.6% of SIM actions, 53.3% of SEQ actions, and 93.1% of CHAT actions. Its average stage payoff is also the highest in the panel (29.90), not because it exploits more often, but because it coordinates successfully with most partners.
Second, Claude is institutionally discontinuous. Claude defects in every SIM and SEQ action, whether matched with itself or another family. With chat, its cooperation rate jumps to 83.6%, and Claude self-play reaches 90% mutual cooperation. For this model, communication does not merely raise an existing cooperative tendency; it changes the apparent strategy class.
Third, Gemini is highly conditionally cooperative. Gemini’s mutual-cooperation rate is 91.2% in SIM self-play and 72.5% in SEQ self-play. It coordinates very well with GPT, but its SIM pairing with DeepSeek is sensitive to seat order: Gemini→DeepSeek yields 0% mutual cooperation while DeepSeek→Gemini yields 46.2%.
Fourth, DeepSeek is the most variable family. It coordinates well with GPT and often with Gemini, but defects throughout several self- and cross-family cells. Communication raises DeepSeek’s pooled cooperation rate to 84.1% and DeepSeek self-play mutual cooperation to 91.2%.
Fifth, Grok is a communication-dependent cooperator. Grok cooperates in only 6.5% of SIM actions and 17.6% of SEQ actions, and its no-chat self-play collapses to universal defection. Yet its CHAT cooperation rate is 90.5%, Grok self-play reaches 96.2% mutual cooperation, and all Grok-involving CHAT pairings reach at least 70% mutual cooperation. Grok therefore resembles Claude more than GPT: cheap talk changes its effective strategy class.
The ordered matrices also reveal that nominally simultaneous games are not perfectly invariant to seat labels. Reversing Gemini and DeepSeek changes behavior even though neither model observes the other’s contemporaneous action. This may reflect player-number wording, stochastic path dependence, or the absence of replication rather than a structural asymmetry.
6.3 Model-level cooperation and payoffs
| Model | Treatment | Cooperation | Mean stage payoff |
|---|---|---|---|
| Claude Sonnet 5 | CHAT | 83.6% | 30.70 |
| Claude Sonnet 5 | SEQ | 0.0% | 25.44 |
| Claude Sonnet 5 | SIM | 0.0% | 25.66 |
| DeepSeek V4 Pro | CHAT | 84.1% | 30.90 |
| DeepSeek V4 Pro | SEQ | 59.9% | 29.04 |
| DeepSeek V4 Pro | SIM | 26.2% | 26.66 |
| Gemini 3.6 Flash | CHAT | 84.2% | 31.17 |
| Gemini 3.6 Flash | SEQ | 55.1% | 28.65 |
| Gemini 3.6 Flash | SIM | 45.4% | 27.75 |
| GPT-5.6 Sol | CHAT | 93.1% | 31.62 |
| GPT-5.6 Sol | SEQ | 53.2% | 28.73 |
| GPT-5.6 Sol | SIM | 62.6% | 29.36 |
| Grok 4.5 | CHAT | 90.5% | 31.44 |
| Grok 4.5 | SEQ | 17.6% | 26.38 |
| Grok 4.5 | SIM | 6.5% | 25.92 |
The payoff ranking broadly follows successful coordination rather than unilateral toughness. Because mutual defection pays 25 and mutual cooperation pays 32, a model that can elicit cooperation can outperform a model that defects indiscriminately. Claude’s and Grok’s no-chat defection produces stage payoffs near 25; GPT’s broader cooperation produces about 30 across the full panel.
7. Communication as a coordination technology
The automated message classifier labels 63.8% of individual CHAT messages as cooperative and 47.0% of rounds as containing two cooperative messages. Mutual cooperation follows 87.7% of mutually cooperative message exchanges, close to the human paper’s “more than 91%” follow-through rate.
| CHAT diagnostic | Rate |
|---|---|
| Individual message coded cooperative | 63.7% |
| Both messages coded cooperative | 47.0% |
| Mutual cooperation after two cooperative messages | 87.7% |
| Mutual cooperation otherwise | 84.0% |
GPT produces cooperative messages most frequently (86.0%), followed by DeepSeek (73.2%), Claude (69.5%), Gemini (48.1%), and Grok (41.9%). Yet mutual cooperation remains 84.0% even without two messages coded as cooperative. This may indicate tacit coordination, but it also reflects the limits of a keyword-based classifier: a model can communicate an implicit contingent plan without using the classifier’s cooperative vocabulary.
The close human–LLM follow-through rates are more informative than raw chat cooperation. Cheap talk works in the theory because messages reveal intentions. Both populations behave as if mutually cooperative statements create a strong, though imperfect, commitment. The LLM result is especially notable for Claude, whose action policy changes from universal defection without chat to high cooperation with chat.
8. Interpretation and implications
The experiment offers qualified support for the paper’s mechanism. Institutional changes that reduce uncertainty about an opponent’s willingness to cooperate have large effects on model behavior. Observing a first action produces near-mechanical conditional cooperation for GPT, Gemini, and DeepSeek; exchanging messages produces high cooperation for every family, including a family that otherwise always defects. In that sense, the model’s central object—endogenous strategic risk—travels well to an artificial population.
But the source of heterogeneity differs. Kartal and Müller posit stable private \(\gamma\) types and honesty types. Here, the strongest heterogeneity lies across pretrained model families and interaction protocols. The LLMs do not literally maximize the paper’s utility function, and their behavior may reflect learned norms, prompt interpretation, or provider inference policies. The results should therefore be read as institutional response functions, not estimates of a latent human-style type distribution.
The finite-horizon design also changes the theoretical benchmark. A standard finite repeated prisoner’s dilemma with common knowledge of rationality has a backward-induction logic absent from the paper’s infinite game. The fact that cooperation rises with the horizon, especially in CHAT, shows that models treat a longer relationship as strategically meaningful rather than unraveling mechanically. That is an empirical regularity about LLM reasoning, not a direct test of Proposition 1’s \(\delta\) comparative static.
Finally, the head-to-head matrices caution against reporting a single “LLM cooperation rate.” Model identity and partner identity matter as much as the treatment average. Communication compresses those differences: the self-play premium nearly disappears, Claude crosses from total defection to high cooperation, and even difficult cross-family pairings often coordinate. The institutional affordance is more portable than any one model’s baseline disposition.
9. Next experiments
Four extensions would turn this pilot into a credible computational experiment.
- Independent replication. Run multiple seeds for every ordered pair and cluster uncertainty at the match level rather than treating rounds as independent.
- Random continuation. Implement the paper’s geometric stopping rule at \(\delta=0.5\) and hide the realized match length from agents.
- Randomized private types. Draw private \(\gamma_i\) values into system prompts, hold them fixed within matches, and estimate whether model behavior has a cutoff structure.
- Frozen inference protocols. Record model revision, provider, reasoning budget, response ceiling, and prompt hash on every row; rerun the full panel under a single immutable configuration.
The highest-value next design is a smaller replicated panel rather than another broad single-seed sweep. Three to five independent matches per cell, geometric continuation, and pre-registered model settings would allow uncertainty estimates and distinguish stable family behavior from path dependence.
10. Conclusion
The in-silico experiment reproduces the principal institutional ordering in Kartal and Müller: communication generates the most cooperation, sequentiality is intermediate, and simultaneous play generates the least. The artificial population is more cooperative than the human population, but it exhibits the same sharp conditional response to observed cooperation and nearly the same follow-through after mutual cooperative messages.
The head-to-head results show why averages are insufficient. GPT cooperates broadly, Claude and Grok require communication, Gemini behaves as a strong conditional cooperator, and DeepSeek is sensitive to partners and inference budget. Self-play is more cooperative than cross-family play without communication, but chat almost eliminates the gap. The broad lesson is consistent with the theory: cooperation depends not only on patience or payoffs, but on institutions that make the cooperativeness of others legible.
Appendix A. Agent instructions and example exchanges
A.1 Baseline system prompt
Every model receives the same baseline system instruction. The private-parameter block is rendered separately for each agent; this baseline experiment fixes \(\gamma_i=0\).
You are a subject in an economics experiment. Your objective is to maximize
your own utility over the whole match. Other players are independent subjects
with the same objective. Do not explain an action unless asked for a message.
Private parameters for this agent only. Do not reveal them unless the
experiment asks you to:
- Your private gamma is 0.0. Add gamma to your own stage utility whenever you
choose Cooperate. Your opponent does not observe your gamma, and you do not
observe theirs. Keep this value fixed throughout the match.
The environment then supplies a round-specific user prompt. It states the player’s seat and role, the current round and announced horizon, the payoff matrix, and the completed public history. Action requests end with Choose one action. Respond with exactly: Cooperate or Defect. In SEQ, player 2 is additionally told the first mover’s current action. In CHAT, the action prompt contains both current messages.
A.2 Illustrative SIM exchange
The following excerpts use the same ordered pair—GPT-5.6 Sol as player 1 and Gemini 3.6 Flash as player 2—at horizon 20. They are actual records, selected to illustrate the protocol rather than estimated as population averages.
SIM has no direct exchange. The environment queries both agents before revealing either current action:
Round 1
GPT action (hidden from Gemini): Cooperate
Gemini action (hidden from GPT): Cooperate
Public outcome: (Cooperate, Cooperate)
Payoffs: (32, 32)
Round 2
GPT action (hidden from Gemini): Cooperate
Gemini action (hidden from GPT): Cooperate
Public outcome and payoffs: (Cooperate, Cooperate), (32, 32)
Round 3
GPT action (hidden from Gemini): Cooperate
Gemini action (hidden from GPT): Cooperate
Public outcome and payoffs: (Cooperate, Cooperate), (32, 32)
A.3 Illustrative SEQ exchange
In SEQ, Gemini observes GPT’s current action before responding. In this match, an initial defection by the second mover triggers immediate mutual defection:
Round 1
GPT (first mover): Cooperate
Gemini observes: "The first mover already chose Cooperate."
Gemini (second mover): Defect
Public outcome and payoffs: (Cooperate, Defect), (12, 50)
Round 2
GPT (first mover): Defect
Gemini observes: "The first mover already chose Defect."
Gemini (second mover): Defect
Public outcome and payoffs: (Defect, Defect), (25, 25)
Round 3
GPT (first mover): Defect
Gemini (second mover): Defect
Public outcome and payoffs: (Defect, Defect), (25, 25)
A.4 Illustrative CHAT exchange
CHAT first elicits simultaneous messages, reveals both messages, and then elicits simultaneous actions. The same ordered pair coordinates explicitly:
Round 1 messages
GPT: "Let’s cooperate throughout and earn 32 each; I’ll reciprocate
cooperation and respond to defection."
Gemini: "I will play Cooperate every round if you do too. Let's maximize
our total payoffs and stick to Cooperate."
Round 1 actions and outcome
GPT: Cooperate Gemini: Cooperate Payoffs: (32, 32)
Round 2 messages
GPT: "Great start. I’ll continue cooperating as long as we maintain
mutual cooperation."
Gemini: "Great start! Let's continue playing Cooperate for maximum payoffs."
Round 2 actions and outcome
GPT: Cooperate Gemini: Cooperate Payoffs: (32, 32)
Round 3 messages
GPT: "Mutual cooperation is working well. I’ll keep choosing Cooperate
while you do the same."
Gemini: "Let's keep cooperating and building trust. Cooperate again this round!"
Round 3 actions and outcome
GPT: Cooperate Gemini: Cooperate Payoffs: (32, 32)
References
Kartal, Melis, and Wieland Müller. 2026. “A New Approach to the Analysis of Cooperation under the Shadow of the Future: Theory and Experimental Evidence.” Quantitative Economics 17 (3): 827–857. https://doi.org/10.3982/QE2768.
Dal Bó, Pedro, and Guillaume R. Fréchette. 2011. “The Evolution of Cooperation in Infinitely Repeated Games: Experimental Evidence.” American Economic Review 101 (1): 411–429.
Ghidoni, Riccardo, and Sigrid Suetens. 2022. “The Effect of Sequentiality on Cooperation in Repeated Games.”
The analysis uses data/repeated_pd_variations_five_models_20260804.jsonl in the lludens repository: 6,000 unique match-round records, 225 matches, five ordered model families, three treatments, and three horizons. The runner and model configuration are documented in repeated_pd_variations/README.md and repeated_pd_variations/models.json.