Skip to content
NAMEFRAMEPricing

The modern city junction capture.

A four-way crossing in a glass and concrete district. Traffic is aligned to the lane it sits in, people are on the pavements and on the crossings, and 150 barrels, bins, cones and bollards stand among them carrying no label at all. Built from three runs under three seeds, so the crowd is arranged three different ways over the same streets. 300 frames, 6,815 labelled instances across 6 classes, graded B · 86.8 out of 100 by the run's own validator, whose findings are listed further down, in full.

Map
DemoMap
Frames in the run
300
Labelled instances
6,815
Resolution
1280×720
Seed
20260909
Validator grade
B · 86.8/100

What it is for

Aerial detection at a signalled junction: crossings, lane traffic and pedestrians mixed together, with five classes rather than three.

How the labels were made

Every label was derived from the engine's own per-instance ID buffer and the camera transform that rendered the frame, not from a model and not by hand. What occludes what is decided by the GPU on the same pass that draws the frame.

Engine
Unreal Engine 5.8
Camera
One camera box over a crossroads, 28 to 46 m, 1280×720
Classes
person, bicycle, bench, umbrella, car, vegetation
Annotation formats
YOLO, COCO, per-instance ID buffers, per-frame metadata JSON
Packaged
300 of 300 frames, across 1 archive
Version
1.0

What the frames look like

One frame from the run, downscaled for the web: the junction with all four crossings in frame, pedestrians on them and traffic waiting. It was picked to show what the scene is, so read it as an illustration rather than as a sample. The archive holds the other 299.

Frame mc_wide_b__plugin_000058.png of the Unreal Engine modern city junction capture: the junction with all four crossings in frame, pedestrians on them and traffic waiting
DemoMap1280×720 per frameOne frame from the run, downscaled for the web

Download the archive.

One archive, holding the whole run.

The same run is published on the Hugging Face Hub, with a dataset viewer over the frames: huggingface.co/datasets/NameFrame/modern-city-unreal-synthetic.

Modern city junction · complete capture

554 MB

All 300 frames of the run: images, YOLO labels, the same labels as COCO, per-instance segmentation, per-frame camera metadata, the capture contract and the run's own report. 6,807 labelled instances. Depth is left out: it is 1.37 GB per capture of float32 nobody training a box detector will open, and the job files to re-capture it are published.

Frames
300
Annotated instances
6,807
Format
ZIP

Counted from the run’s own labels: 300 of 300 requested frames matched a label file in the export.

  • images/ rendered RGB frames, unmodified PNG
  • segmentation/ per-instance ID buffers, one colour per instance
  • labels/ YOLO boxes, one text file per frame
  • annotations/instances.json the same labels in COCO format
  • metadata/ per-frame camera, actor and spawn-manifest records
  • reports/ this capture's own report and validation output
  • capture.json the run contract: classes, camera policy, seed
  • data.yaml class names, ready for a YOLO trainer
Download 554 MB

SHA-256c588ededf66edbdd16bfb20f6fde2ca26b09212b4a5ec8d9aa4828e7ba922734

Known limitations

The validator graded this run B · 86.8/100, and its report ships inside every archive. Here is what it and we know is wrong with the data.

  • The validator took 7.87 points off for class-imbalance: rarest/commonest class ratio 0.016. It is one of the reasons the run scored 86.8 rather than 100.
  • The validator took 5 points off for near-duplicate-images: visually near-identical images in one split (dHash <= 4) (71 of 300). It is one of the reasons the run scored 86.8 rather than 100.
  • The validator took 0.23 points off for blurry-images: variance-of-Laplacian < 100.0 (7 of 300). It is one of the reasons the run scored 86.8 rather than 100.
  • The validator took 0.13 points off for underexposed-images: mean luminance < 35.0 (5 of 300). It is one of the reasons the run scored 86.8 rather than 100.
  • The capture checked itself for people it could see and did not label: 3/2709 visible (292 frames engine, 8 depth) people unlabelled (0.11%, 300/300 frames). The check ships inside the pack.
  • One map, one seed and one camera policy. A model trained on this capture alone has seen one place; the eight captures are published together for that reason.
  • The archive carries COCO boxes rather than COCO polygons. The per-instance ID buffers ship alongside them and are the mask ground truth.
  • A detector trained on all eight of these captures and nothing else reached 0.350 recall on real drone footage. Synthetic data alone did not close that gap, and this archive is published so the number can be argued with rather than believed.

The checks behind those numbers, and how a run is graded, are on the dataset quality page, and the validator itself is described under dataset validation.

Terms, plainly

Free to download, read, train on and benchmark against. Publish whatever results you get, including bad ones. What you may not do is repackage the pack and sell it as your own dataset.

The images depict environment and prop content licensed for use inside Unreal Engine projects. Redistribution rights for that underlying content are not granted with these packs, so check the source licence before publishing derivatives. The full text ships as TERMS.txt inside each archive.

Citation

If you publish anything you got out of this data, this is how to point at the exact run it came from. The seed and the map are the part that matters: they are what makes the run reproducible.

NameFrame. "Modern city junction: synthetic computer-vision dataset." Version 1.0, generated with Unreal Engine 5.8 on DemoMap, seed 20260909. https://getnameframe.com/datasets/modern-city

Need this scene, but yours?

This capture came out of one NameFrame run: a map, a class list, a camera policy and a seed. Change any of them and you get a different dataset with the same ground truth guarantees. That is what the generator is for.