# Florence-2 boxes to game-reference crops: offline evidence

This is a small, inspectable experiment in turning pixel bounding boxes into
integer image crops. The two scenes are original synthetic geometric drawings.
Their object labels and boxes were deliberately specified in code. They are
not Florence-2 predictions, game-engine captures, or Goblin3D product outputs.

No inference, model download, remote source execution, paid call, or network
request occurs in these scripts. They do not measure detection accuracy,
segmentation, background removal, 3D quality, or game-asset readiness.

## Reproduce

Python 3.10 or newer, standard library only. No installation step is needed.
Run these commands from this directory:

```sh
python3 -m unittest -v test_crop_boxes
python3 build_evidence.py --check
python3 crop_boxes.py < artifacts/fixtures.json
```

The recorded run passed **72 named unittest tests** (with additional subcases).
The byte-reproduction check matched **15 artifact files**. See `test-results.txt`
for actual test output. Tested with CPython 3.12.14 and zlib 1.3.2. Byte-identical
PNG compression across other zlib versions is not promised; geometry and RGB
pixels are the relevant invariant. `--check` intentionally checks exact bytes.

To generate from scratch, copy the Python source files into a new directory that
has no `artifacts` subdirectory, then run `python3 build_evidence.py`. The builder
writes only a fixed set of filenames in `artifacts/`; it refuses an existing
output directory. `--check` regenerates in memory and does not modify outputs.
The crop-validation CLI reads a bounded JSON request from stdin and returns
JSON on stdout. It never opens images, creates paths from labels, or writes
output files. Redirect stdout yourself if you want to save its result.

## The coordinate contract

1. Decode the actual raster and normalize its EXIF orientation **before** assigning
   a frame identifier or obtaining boxes. Pass that same upright image into the
   detector. Record its actual `(width, height)` after normalization. This package
   does not decode EXIF or perform orientation changes; its PNG fixtures contain
   no orientation metadata. A true flag is a caller assertion, not proof.
2. An already post-processed `<OD>` object contains `bboxes` paired with `labels`.
   This adapter expects pixel `[x0, y0, x1, y1]`, with top-left origin, x increasing
   rightward and y increasing downward. Raw location tokens, normalized 0–1
   coordinates, OCR quadrilaterals, and polygons are not that input format.
3. Require four finite Python int/float values, excluding booleans; require
   `x0 < x1` and `y0 < y1`. Reject a box with no positive-area overlap with the
   raster **before padding**. Never sort reversed edges to conceal an error.
4. Add optional nonnegative symmetric context padding in destination pixels.
   Then floor left/top, ceil right/bottom, and clip all four edges to the raster.
   Output is half-open `[left, right) × [top, bottom)`, so crop dimensions are
   `(right-left, bottom-top)`. Padding expands the selected context; it does not
   synthesize border pixels or guarantee a fixed margin at an image edge.
5. Use the same declared frame for the boxes and raster. Different frame IDs
   are rejected. A false matching ID, swapped dimensions, or mislabeled units
   cannot be detected reliably from four numbers; callers must preserve provenance.
   For example, coordinates below 1 can also be legitimate subpixel boxes.
6. A resize remap is optional and narrowly supported. Both dimensions must have
   exactly the same aspect ratio, and the caller must explicitly assert an aligned
   full-raster uniform resize. Use a distinct frame ID when dimensions change.
   Scale all four edges before padding and rounding. Same aspect ratio alone
   does not establish alignment or identical content. Letterboxing, cropping,
   mirroring, rotation, offsets, and changed content require their real transform;
   this function does not guess or silently correct them.

The dimension bound of 1,000,000 in the validator is an arithmetic input limit,
not permission to allocate huge images. The demonstration renderer caps raster
axes at 4096. JSON requests are capped at 1,000,000 characters and 10,000 records.
These are narrow educational utilities, not a hardened image-serving service.

## What was measured

The 960 × 540 wide fixture has 4 accepted and 5 rejected records. With 6 pixels
of context padding:

| Synthetic item | Integer crop [left, top, right, bottom] | Size, width × height |
| --- | --- | --- |
| Sword | [42, 260, 447, 360] | 405 × 100 |
| Staff | [522, 73, 608, 432] | 86 × 359 |
| Shield | [689, 108, 888, 351] | 199 × 243 |
| Border potion | [0, 378, 78, 540] | 78 × 162 |

The potion is intentionally cut off in the source raster: clipping preserves
only existing pixels and cannot reconstruct the missing part. Rejections cover
fully outside, zero-width, reversed-x, boolean, and text-valued coordinates.
Additional tests cover nonfinite values, all outside directions and edge-touch
cases, fractional padding, overflow, malformed structure, and exact pixel copying.

The independent tall fixture is 360 × 720. Its staff box
`[220.2, 55.7, 273.1, 658.2]`, padded by 6 pixels, becomes
`[214, 49, 280, 665]`, a 66 × 616 crop. The asymmetric frames help expose an
accidental width/height swap.

For the wide scene resized to 480 × 270, the sword's original box
`[48.25, 266.5, 440.75, 353.25]` maps to
`[24.125, 133.25, 220.375, 176.625]`. Without padding, rounding gives
`[24, 133, 221, 177]`, a 197 × 44 crop. The deliberately incorrect use of the
original coordinates on the smaller raster gives only 393 × 4 pixels. The
wrong-frame diagram intentionally supplies a false matching frame ID to expose
that numeric bounds checks alone cannot certify coordinate provenance. With the
true differing ID, the validator rejects the crop.

## Source boundary

The public Microsoft processor was inspected at the immutable repository revision
`5ca5edf5bd017b9919c05d08aebef5e4c7ac3bac`:

- [Pinned processor](https://huggingface.co/microsoft/Florence-2-base/blob/5ca5edf5bd017b9919c05d08aebef5e4c7ac3bac/processing_florence2.py)
- [Public model card: object detection](https://huggingface.co/microsoft/Florence-2-base#object-detection)
- [Pinned repository tree](https://huggingface.co/microsoft/Florence-2-base/tree/5ca5edf5bd017b9919c05d08aebef5e4c7ac3bac)

The model card shows OD's bbox/label structure and passes `(image.width,
image.height)` to post-processing. In the pinned processor, `BoxQuantizer`
interprets axes as width then height and the default box quantization uses 1000
bins per axis. Its dequantization uses a bin's center. The isolated
`florence_1000_bin_centers` function illustrates that arithmetic only:
`(index + 0.5) × axis_size / 1000`. For `[0, 0, 999, 999]` in a 960 × 540 frame,
it yields `[0.48, 0.27, 959.52, 539.73]`. This function is not the processor, does
not parse generated text, and does not establish tensor-level numerical identity.
Already post-processed pixel bboxes must not be dequantized again.

The optional upstream `scores` field is ignored by this adapter; these tests
make no confidence calibration or detection-quality claims. The integer rounding,
clipping, padding, frame assertion, and rejection rules here are our explicit
crop policy, not claims about a Florence-2 crop API.

## Artifacts and metadata

- `artifacts/scene-wide.png`: 960 × 540 unannotated original scene
- `artifacts/scene-tall.png`: 360 × 720 independent tall scene
- `artifacts/scene-wide-boxes.png`: padded integer crop boundaries in mint
- `artifacts/crop-00-sword.png` through `crop-04-tall-staff.png`: exact RGB crops
- `artifacts/crop-contact-sheet.png`: 1440 × 900 overview; fitted card previews
- `artifacts/resize-half.png`: 480 × 270 nearest-neighbor display-resize fixture
- `artifacts/resize-correct-crop.png`, `resize-wrong-crop.png`: contrasting results
- `artifacts/coordinate-frame-comparison.png`: 1560 × 675 frame demonstration
- `artifacts/fixtures.json`: reproducible synthetic post-processed-OD-shaped input
- `artifacts/report.json`: geometry, evidence limits, image sizes, PNG hashes

All PNGs are lossless RGB, 8 bits per channel, with only IHDR, IDAT, and IEND
chunks. They have no EXIF, ICC profile, text, author, tool, time, or location
metadata. Display browsers may apply their normal untagged-RGB color handling.
Source images and crops have no interpolation; only the contact-sheet previews
and half-size comparison use documented nearest-neighbor resizing. The contact
sheet's card backgrounds are presentation layout, not pixels added to crop files.

## License and provenance

Original code, diagrams, scenes, JSON, and documentation are released under the
included MIT license. No externally authored image or upstream source file is
bundled. The referenced model repository advertises MIT, while the pinned
`processing_florence2.py` file itself has an Apache-2.0 header. These are separate
upstream licensing facts; the MIT license here does not relicense that source.
No model or upstream source file was executed. `provenance.json` records the public sources,
runtime and verification counts. No operating-system wall-clock timestamp is used
as evidence of when this experiment ran.
