How corner is a corner case? Percentile control for highway scenario generation

Jiaxi Liu, Hang Zhou, Hangyu Li, Yifan Wang, Keke Long, Chengyuan Ma, Bin Ran, Xiaopeng Li*

Department of Civil and Environmental Engineering, University of Wisconsin–Madison

*Corresponding author: xli2485@wisc.edu

Same history, three requests, filmed from above. An IDM+MOBIL planner (orange) drives through futures that our method generates at the 10th, 50th and 90th percentiles of this history's own risk distribution, while the other vehicles replay those futures without reacting. Blue marks the vehicle that sets the planner's minimum post-encroachment time (PET). At p = 0.1 no vehicle in view changes lanes, and the blue vehicle is a leader more than 80 m ahead, outside the view (arrow). At p = 0.5 the vehicle in the next lane falls back and is just reaching the lane line, 29 m behind the planner, when the window ends. At p = 0.9 it moves into the planner's lane 14 m behind it. The minimum PET falls from 2.74 s to 0.77 s and 0.30 s.

Abstract

Generating corner-case scenarios with appropriate adversity in a simulation environment is critical for testing an autonomous vehicle (AV) software stack’s safety performance before deployment. Existing autonomous-driving scenario generators can enforce specific behavior, adversity, or feasibility conditions, but they provide limited control over how extreme a generated scenario is relative to plausible futures in the same traffic context. This study represents the adversity of a generated scenario as its percentile in the conditional distribution of future risk given the observed history. This view supports calibrated answers to two questions: how “corner” a generated corner-case scenario is and how its “cornerness” can be fine-tuned.

To this end, we formulate history-conditioned risk-percentile requests and learn a reference risk distribution that maps each requested percentile to a physical risk target. We then use a percentile-conditioned joint diffusion model with sampling-time risk guidance to generate multi-agent futures, together with a reference-based criterion for evaluating percentile realization.

Experiments use the minimum post-encroachment time (PET) between the ego and its surrounding vehicles as the risk surrogate on highD. On the primary evaluation set, our method realizes 1,422 of 1,440 requests within a 0.05 percentile tolerance (98.75%), with mean percentile error 0.00673 and PET-target error 0.00991 seconds. The resulting interface connects context-relative risk specification, physical realization, and evaluation through a common risk scale.

Why adversity percentiles?

The same one-second PET can be routine in one traffic context and rare in another. A physical threshold alone therefore says little about how extreme a scenario is for its situation. We index a corner case by its percentile among the natural futures of the observed history, so a request such as p = 0.9 means the same degree of adversity in sparse and in dense traffic.

Existing scenario-generation interfaces take tokens, language prompts or constraints as control inputs.
(a) Existing interfaces use tokens, language prompts, or constraints and leave adversity percentiles unquantified.
Our interface takes a target adversity percentile p as the control input.
(b) Our view: corner cases are indexed by a target adversity percentile p.

Method

Three-layer workflow of data preprocessing, training and testing with the reference risk distribution and the percentile-conditioned joint generator.

Three-layer workflow: data preprocessing, training and testing. The percentile predictor denotes the reference risk distribution, trained by CRPS on observed scene-risk measurements. The request enters the joint generator directly and guides its sampling path, together with road and background-separation constraints.

  1. Reference. A risk distribution learned from natural highD trajectories maps each requested percentile p to a physical PET target for the observed history.
  2. Generation. A percentile-conditioned joint diffusion model generates the futures of all vehicles, and sampling-time risk guidance moves the outcome toward the target while road and background-separation constraints keep the scene feasible.
  3. Evaluation. A reference-based criterion checks whether the generated future lands at p, including outcomes at probability masses such as the PET cap.

The reference and the generator are one implementation of the request. Any calibrated conditional distribution can serve as the reference, and any generator that can be steered toward a physical target can realize the request.

Results

98.75%of 1,440 requests realized within 0.05 of the requested percentile, with mean percentile error 0.00673
13–17%of the requests met by the compared generators, even with a percentile input
98.61%of time-to-collision requests in car following, met by the same generator without retraining
Heatmaps of requested against realized percentile for our method, P-CVAE, Standard P-diffusion, RADE and natural priors.
Where the outputs land. Each column shows, for one requested percentile, the share of outputs whose realized percentile falls in each bin. The outputs of our method concentrate at the request, while those of P-CVAE, Standard P-diffusion, RADE and natural priors spread over the whole scale.
Share of requests met as a function of the percentile tolerance, and when the closest of K samples of the same history is selected.
Precision across tolerances and selection among samples. (a) Share of the 1,440 requests realized within a percentile tolerance. (b) Share realized within 0.05 when the sample whose percentile is closest to the request is chosen among K samples of the same history. Even among 15 samples, the CTG++ prior meets 34.4% of the requests and the STRIVE prior 24.6%.

Comparison with external generators

External scenario generators learn natural traffic but take no percentile input. On one car-following history, we even give two of them an advantage: from the 15 samples of CTG++ and STRIVE, we pick after generation the one whose PET comes closest to each target. TrafficGen gives one output per history. Our method still lands closer to every target, and across all 96 test histories the external priors meet only 13–16% of the requests within the 0.05 tolerance.

Our method at three requested percentiles and the TrafficGen, CTG++ and STRIVE samples closest to each target on the same car-following history, with the PET of all samples.
The same car-following history with 10 vehicles. (a) Our method at p = 0.1, 0.5 and 0.9. (b) TrafficGen: the top-1 output, generated without a percentile input. (c) CTG++ unguided prior and (d) STRIVE traffic prior: for each target, the sample whose PET is closest to it among 15 samples, selected after generation. (e) PET of all samples, with dotted lines at the targets. Orange marks the ego, blue the vehicle that sets its minimum PET and gray the other vehicles.

Planner test

A percentile request is useful for AV testing when it translates into graded difficulty for the system under test. We replace the ego of each generated scenario with an IDM+MOBIL planner, and the surrounding vehicles replay their generated futures without reacting. The videos are rendered in MetaDrive from the exact trajectories of the paper's planner test, in real time at 25 frames per second. Each case shows one history under three requests, and the first four are the histories of Figs. 11 and 12 of the paper. In both views the planner is orange, the vehicle that reaches the minimum PET with it is blue and the other vehicles are gray. In the top-down view, a red cross marks the point both of them pass.

Follower closing in behind the planner

7 vehicles · 3 lanes · Fig. 11(a) of the paper

At p = 0.9, a follower 50 m behind in the planner's lane accelerates from 34.0 to 45.8 m/s and closes in.

Planner minimum PET, bars from 0 to 3 s
p = 0.11.37 s
p = 0.50.53 s
p = 0.90.20 s
Chase view
Top-down view

BibTeX

@article{liu2026corner,
  title   = {How corner is a corner case? Percentile control for
             highway scenario generation},
  author  = {Liu, Jiaxi and Zhou, Hang and Li, Hangyu and Wang, Yifan and
             Long, Keke and Ma, Chengyuan and Ran, Bin and Li, Xiaopeng},
  journal = {Transportation Research Part C: Emerging Technologies},
  note    = {Under review},
  year    = {2026}
}