PG-DPO

Pontryagin-guided control recovery

A common computational architecture: simulate a continuation policy, estimate adapted adjoints by BPTT, then enforce the problem-specific local Pontryagin condition in action space.

Schematic educational visualization. Curves and diagnostics illustrate mechanisms; they are not reproduced experiment logs.
Example
Stage I Rollout warm-up
Stage II-A Adjoint estimation
Stage II-B Local recovery

The papers group II-A and II-B together as Stage II; the playground separates them to make the mechanism visible.

Figure 1

Rollout warm-up

simulation

Figure 2

Adapted adjoints

BPTT

Figure 3

Local recovery

Hamiltonian

Research scope