Skip to content

Latest commit

 

History

32 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

English | 简体中文

👋 Hi, everyone! DanKS is a GuanDan AI project initiated by the Kingsoft AI Product Center.

CI Release AtomGit Python 3.10+ Apache-2.0 license

Kingsoft AI Product Center

DanKS: State-of-the-art GuanDan AI

Three complete generations of code

Play online · Quick start · Architecture · Generations · Training · CardKS paper hub

Meet DanKS—the state-of-the-art AI built to master four-player, partnership-based GuanDan. This single repository reveals its complete three-generation evolution: from structural retrieval and learned candidate selection to a memory-aware policy trained with PPO—all powered by a shared 108-card GuanDan rules engine.

Online demo

DanKS promotional hero with the Kingsoft AI Product Center logo and online GuanDan table

▶ Challenge DanKS in your browser
No local setup · one human seat and three bot seats · Chinese and English interface

Watch a short gameplay preview

Short animated preview of the CardKS online GuanDan demo
View the full-resolution live table · Download the 1280 × 640 social preview

Quick start

The shortest runnable path uses V3 on CPU. Commands below assume Python 3.11 and a POSIX shell:

git clone https://github.com/Calix-L/DanKS.git
cd DanKS
python3.11 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e versions/v3
python -m pip install torch==2.8.0 --index-url https://download.pytorch.org/whl/cpu
python examples/retrieval_quickstart.py --version v3
python examples/v3_model_smoke.py

Use .venv\Scripts\Activate.ps1 on Windows PowerShell. CUDA, Ascend NPU, V1/V2, native-kernel, and development setups are documented in Installation reference.

Overall architecture

Overall DanKS pipeline shared by the three versions, from GuanDan information state and structured candidate retrieval to actor-critic scoring and PPO self-play

DanKS turns a large, structured action space into a compact policy decision:

  1. Encode the information state. The policy receives the visible hand, public action history, legal actions, and seat-aware game context.
  2. Retrieve structured candidates. Budgeted decomposition search produces representative plays and summarizes their length, pairs, sequences, suits, gaps, and remaining-hand structure.
  3. Score a bounded Top-K set. A shared encoder combines state, candidate, and structural features; the actor ranks valid candidates while the critic estimates state value.
  4. Learn from self-play. Trajectories provide GAE advantages for clipped PPO updates, improving the selector without expanding the inference-time candidate budget.

This shared architecture spans all three generations of DanKS, covering the complete path from information state and structured candidate retrieval to policy/value estimation and self-play optimization. Each generation advances the features, retrieval strategy, and policy implementation within this framework.

Why DanKS?

  • State-of-the-art playing strength — DanKS achieves leading results against strong learning-based and rule-based GuanDan baselines under the complete promotion-match protocol; see the CardKS main results.
  • A clear three-generation codebase — V1, V2, and V3 expose the full technical progression, making each major algorithmic advance easy to read, run, and compare.
  • The complete pipeline is included — the repository covers the GuanDan rules engine, legal-action generation, structured retrieval, state and candidate features, policy/value models, PPO training, checkpoint handling, native acceleration, and runnable inference examples.

Generations

Version Main idea What it adds Entry point
V1 Structural retrieval Candidate scoring and a NumPy selector ranker.py
V2 Learned selection Broader action generation and an ONNX selector action_generator.py
V3 Memory-aware policy learning Card memory, candidate coverage, recall, team belief, and PPO model.py

V1, V2, and V3 are available as separate packages. Give each generation its own environment to keep the DanKS import, feature schema, and checkpoint format aligned.

Why delayed outcomes matter

Three candidate actions from the same GuanDan state leading to different delayed structural outcomes

A move that looks cheap now can destroy the only useful combination left in the hand; spending a powerful card can preserve structure and create a cleaner future exit. DanKS separates the responsibilities needed to learn that distinction:

  • Retrieval organizes the combinatorial action space into a strategically varied candidate set.
  • Structure features expose what each candidate consumes, preserves, or leaves behind.
  • The actor scores the legal candidates available in the current state.
  • The critic and GAE assign credit from later trajectory outcomes, allowing PPO to favor actions whose value appears several decisions later.

The illustration captures the central idea behind long-horizon credit assignment: V3 scores retrieved candidates directly and learns their long-term value from subsequent trajectories.

Repository layout

DanKS/
├── assets/             # brand, online demo, architecture, and decision figures
├── versions/
│   ├── v1/DanKS/       # retrieval + NumPy selector
│   ├── v2/DanKS/       # retrieval + ONNX selector
│   └── v3/DanKS/       # retrieval + neural policy + PPO
├── guandan/engine/     # shared Python rules engine
├── examples/           # executable engine, retrieval, and model smoke runs
├── tests/              # repository and engine checks
├── README.zh-CN.md     # complete Simplified Chinese guide
├── pyproject.toml
├── LICENSE
└── NOTICE

Installation reference

The shared engine and V3 support Python 3.10 and newer; V1 and V2 support Python 3.11 and newer. All commands below run from the repository root. A dedicated virtual environment for each generation keeps the DanKS namespace aligned with its features and model format.

Choose a package

Goal Install command Notes
Shared rules engine and tests python -m pip install -e '.[dev]' Rules engine and repository test suite.
V1 · structural retrieval python -m pip install -e versions/v1 NumPy selector; Python 3.11+.
V2 · learned selection python -m pip install -e versions/v2 ONNX selector; Python 3.11+.
V3 · PPO policy python -m pip install -e versions/v3 Install one PyTorch build below.

Select one V3 PyTorch build

Target Command
Linux / Windows CPU python -m pip install torch==2.8.0 --index-url https://download.pytorch.org/whl/cpu
NVIDIA CUDA 12.8 python -m pip install torch==2.8.0 --index-url https://download.pytorch.org/whl/cu128
macOS CPU python -m pip install torch==2.8.0

Use the official PyTorch installation matrix when your platform requires a different wheel. Check the resulting installation with python examples/v3_model_smoke.py and inspect all learner options with python -m DanKS.training.train_ppo --help.

Ascend NPU setup

The Ascend runtime works with matching host drivers and CANN releases. Install the matching CANN release, followed by the vendor-provided PyTorch and torch_npu wheels. A validated combination is recorded in requirements-training-npu.txt.

source /usr/local/Ascend/cann/set_env.sh
python3.10 -m venv --system-site-packages .venv-v3-npu
source .venv-v3-npu/bin/activate
python -m pip install -e versions/v3
python -m pip install --no-deps \
  /path/to/torch-2.7.1+cpu-cp310-cp310-manylinux_2_28_x86_64.whl \
  /path/to/torch_npu-2.7.1.post2-cp310-cp310-manylinux_2_28_x86_64.whl
export TORCH_DEVICE_BACKEND_AUTOLOAD=0
python -m DanKS.training.train_ppo --help

Keep this virtual environment dedicated to Ascend NPU. For other driver, CANN, architecture, or Python combinations, select the corresponding vendor wheels.

Validated configurations
Target System Python Framework Key packages
CI and shared engine Linux 3.10, 3.12 pytest 7+
V1 CPU 3.11+ NumPy selector NumPy 2.4.6
V2 CPU 3.11+ ONNX selector NumPy 2.4.6, ONNX Runtime 1.27.0
V3 NVIDIA server H100, driver 575.57.08 3.11.14 PyTorch 2.8.0 + CUDA 12.8 NumPy 2.4.6, pybind11 3.0.4
V3 Ascend server Ubuntu 22.04.5, 910B2C, driver 24.1.0, CANN 8.5.0 3.10.12 PyTorch 2.7.1 + torch_npu 2.7.1.post2 NumPy 1.26.0, pybind11 3.0.4

These are known-good reference configurations; DanKS also runs on other compatible environments.

Optional V3 C++ acceleration

The optimized retrieval kernels support Linux and macOS with a C++17 compiler, Python development headers, and pybind11; Windows automatically selects the Python implementation. Install the platform toolchain once:

# Ubuntu/Debian
sudo apt-get update && sudo apt-get install -y build-essential python3-dev

# macOS (run once)
xcode-select --install

Then build and verify both kernels with one command in the active V3 environment:

danks-build-native

The command locates the installed V3 source tree automatically and finishes with cover=True, actor=True. Linux builds enable host-specific compiler optimization; macOS delegates architecture selection to the Python toolchain and supports universal2 builds. Run it again after changing Python versions or CPU architecture. Windows automatically selects the Python implementation.

Development checks

python -m pip install -e '.[dev]'
python -m pytest -q

Run the examples

The examples cover the rules engine, structural retrieval, a full network forward pass, and a PPO update, all directly runnable from source:

# Shared rules engine; available from the base environment.
python examples/engine_quickstart.py

# Structural retrieval; run inside a matching V1, V2, or V3 environment.
python examples/retrieval_quickstart.py --version v3

# Full V3 network forward pass; run inside a V3 environment with PyTorch.
python examples/v3_model_smoke.py

# One synthetic optimizer update through the V3 PPO learner.
python examples/v3_ppo_smoke.py

Each command includes self-checking assertions for a quick confirmation that the environment and code path are working.

Shared game engine

The engine can be used independently of the AI generations:

from guandan import Environment

game = Environment(first_player=0)
for seat in range(4):
    game.add_player(f"player-{seat}", seat)

messages = game.start()
assert all(len(player.hand_cards) == 27 for player in game.players)

The public API also exports Move and Moves for move representation and legal-action generation.

Train V3 with PPO

After activating and verifying a V3 environment, select the rollout and checkpoint output paths:

python -m DanKS.training.train_ppo \
  --rollout /path/to/rollout.npz \
  --output /path/to/checkpoint.pt \
  --device auto

The learner expects rollout arrays for state, candidates, masks, history, actions, behavior log-probabilities, advantages, and returns. Run the entry point with --help for optimization, evaluation, accelerator, and initialization options.

The V3 training implementation lives in versions/v3/DanKS/training and includes:

  • model and feature definitions;
  • PPO objectives and tactical resampling;
  • checkpoint and optimizer-state handling;
  • persistent learner transport;
  • recall and team-belief auxiliary paths;
  • CPU, CUDA, and NPU-aware accelerator helpers.

Contributing

Bug fixes, tests, portability improvements, and algorithmic advances are welcome. See the contribution guide to get started.

License

DanKS is available under the Apache License 2.0. Third-party dependencies retain their respective licenses; see NOTICE.

About

RL‑Empowered Small‑Scale Competitive Guandan Agent

Topics

Resources

Contributing

Stars

171 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages