MiniLM Game-Item Search: Similarity Is Not a Game Rule

A reproducible eight-item MiniLM experiment shows useful matches, a near tie, and failures. Learn where semantic retrieval ends and game rules begin.

colinkoko6 min read

Key takeaways

  • Use sentence embeddings to suggest candidate items, then apply exact IDs and structured game rules.
  • The fixed test includes both useful paraphrase matches and failures; a top-ranked result is not proof that an item has the requested ability.
  • Preserve model revision, dtype, pooling, normalization, input text, and batching when reproducing rankings.

Use embeddings to find candidates, not to decide what an item can do

A player may ask for something that helps them cross a chasm without remembering the name “climbing rope.” Sentence embeddings offer a way to rank item descriptions by meaning. They do not establish whether an item is usable, owned, unlocked, or capable of satisfying a game rule. Keep those decisions in explicit data and code.[3]

This guide tests that boundary with eight invented inventory items and seven fixed English queries. The useful result is not a claim that semantic search always works: the actual run contains a near tie, an unavailable ability, and a negation failure. All inputs, vectors, rankings, and reproduction code are linked below.

A floating magnifying glass and glowing speech bubble above a lantern, rope, potion, and key in a fantasy archive.
Concept illustration of searching a game-item catalog by meaning. It is not a product screenshot, model output from this experiment, or a depiction of measured search accuracy.

The Hugging Face resource and the exact setup

Xenova/all-MiniLM-L6-v2 packages ONNX weights for Transformers.js. Its card demonstrates mean pooling and normalization for sentence embeddings. The original Sentence Transformers card describes a 384-dimensional representation and lists semantic search among its intended uses. Both cards identify Apache-2.0 licensing. This is an established resource, checked on October 3, 2026, not a newly announced model.[1][2]

The experiment used Transformers.js 3.8.1, ONNX Runtime Node 1.21.0, q8 weights, and the CPU on Linux x64 under Node 24.19.0. The model revision was 751bff37182d3f1213fa05d7196b954e230abad9. The downloaded model_quantized.onnx file was 22,972,370 bytes; that is one file’s size, not total download size or memory usage. The report records all four downloaded files and their SHA-256 hashes.

The eight descriptions and seven queries were encoded together in that order. Mean pooling and L2 normalization yielded fifteen 384-value vectors. The ranking script computes their dot products; for normalized vectors this corresponds to cosine similarity. The semantic-search documentation describes this normalized-vector approach.[3]

This execution used the Node binding, not a browser. Microsoft documents the Node package and its platform backends separately. Transformers.js also documents browser execution through WASM or WebGPU, but those paths were not exercised here. Do not reuse these results as browser timing or compatibility evidence.[5][4]

What the fixed queries actually returned

The descriptions cover a lantern, rope, healing potion, key, shield, frost wand, apple, and fishing rod. Only each item’s descriptive sentence was embedded: IDs and display names remained separate metadata. The fixture was written before inference and was not rewritten to improve the outcome. Two consecutive initial runs produced byte-identical rankings and vectors.

QueryFirst candidateScoreInterpretation
something to help me see in a caveIron key0.321754Lantern was second, almost tied
recover after being injuredHealing potion0.342371Matches the declared recovery use
get across a deep chasmClimbing rope0.317745Matches the declared crossing use
open a locked doorIron key0.473133Matches the declared unlocking use
breathe underwaterHealing potion0.131016No item grants underwater breathing
LAN-001Wooden shield0.155848Exact ID was absent from embedded text
a tool that is not a light sourceBrass lantern0.506932Top result contradicts the negation
Observed top result from this one fixed q8 CPU fixture; scores are similarity values, not probabilities

For the cave query, the key scored 0.321754 and the lantern 0.321741, rounded to six decimals. Their unrounded difference was only about 0.000013. A UI that displays three decimal places would show 0.322 for both while still choosing one first. The ranking is reproducible in this setup, but that tiny margin is a poor basis for automatically choosing an item.

The injury, chasm, and locked-door queries each put the intended item first. That demonstrates useful matches in this fixture; it does not establish an accuracy rate for real player language. The seven prompts were deliberately selected, the catalog is tiny, and no representative evaluation set or keyword-search baseline was measured.

Three boundaries a game should enforce explicitly

First, nearest-neighbor search always has a nearest candidate when the catalog is nonempty. “Breathe underwater” returned the healing potion even though its description only promises recovery. A ranker cannot add a missing capability. Use a separate capability check before exposing an actionable result, and allow a clear “no suitable item” state.

Second, an inventory code is a key, not a semantic question. LAN-001 ranked the wooden shield first in this description-only experiment. Route a recognized exact ID to a deterministic lookup before semantic retrieval. The failure does not show that the model forgot an ID: that ID was never part of the embedded item text.

Third, a negative phrase is not a reliable exclusion filter. The request for a tool that is not a light source ranked the lantern first at 0.506932. That score was higher than several useful matches above. A single score threshold would not solve this example. If the player selects “exclude light sources,” apply that constraint to a structured field instead of hoping a vector encodes the rule.

A practical design for an inventory or asset-browser prototype

An implementation worth trying is: exact ID lookup, explicit filters, semantic candidate ranking, then a visible shortlist with descriptions. Keep inventory ownership, quest state, allowed equipment slots, and permissions outside the embedding score. This is a proposed design, not a game integration tested for this article.

For a 3D asset browser, descriptions can index design intent such as “a worn brass lamp for a cave scene.” Store the actual asset ID, file location, rights record, and validated technical properties separately. Retrieval can help someone find a prop; it cannot certify that the mesh is rigged, its license fits the project, or it meets a runtime budget.

Before shipping, collect real permitted queries and label acceptable results, ambiguous requests, and no-match cases. Compare against the exact-name and tag search you already have. Keep failures in the evaluation set. Test different descriptions, model choices, and precision settings as new experiments rather than silently replacing the published fixture.

Reproduce the result before adapting it

Save catalog.json and run-search.mjs from the evidence links into the same directory. In a separate Node project, install @huggingface/transformers@3.8.1, then run node run-search.mjs. On Linux CPU-only setups, ONNX Runtime supports skipping its optional CUDA download with ONNXRUNTIME_NODE_INSTALL_CUDA=skip during installation. The script fetches the pinned public files, checks the fixture hash, writes embeddings.json and search-results.json, and disposes the pipeline. No API key or remote inference service is used.[6]

To inspect the arithmetic without loading a model, use the published vectors and recompute the dot product between a query vector and each of the first eight vectors. Preserve the full scores before rounding. Re-running inference with different inputs, batching, runtime, or precision can change the numbers; record those differences instead of treating a new run as identical.

The original model card says inputs longer than 256 word pieces are truncated by default and describes an English model. This short-English fixture does not establish behavior for long lore entries, Korean queries, multilingual search, or arbitrary negation. Evaluate those separately. The downloadable fixture contains original fictional descriptions; the package redistributes neither model weights nor third-party game assets.[2]

Evidence used

Sources

  1. 1.Pinned MiniLM ONNX model card — Hugging Face. Accessed 2026-10-03.
  2. 2.MiniLM model card: intended uses, length limit, and license — Sentence Transformers. Accessed 2026-10-03.
  3. 3.Sentence Transformers semantic-search documentation — Sentence Transformers. Accessed 2026-10-03.
  4. 4.Transformers.js 3.8.1 documentation — Hugging Face. Accessed 2026-10-03.
  5. 5.ONNX Runtime Node.js binding and supported backends — Microsoft. Accessed 2026-10-03.
  6. 6.ONNX Runtime 1.21.0 installation script: optional CUDA download — Microsoft. Accessed 2026-10-03.

Keep asset creation and asset validation connected

Explore Goblin3D’s public features, then maintain clear descriptions and verification records for the assets you use.

Explore Goblin3D features

Sources, product facts, and original evidence were checked before publication.

Share your feedback

Sign in to share feedback with the Goblin3D team.

Sign in

support@goblin3d.ai