Key takeaways
- Use sentence embeddings to suggest candidate items, then apply exact IDs and structured game rules.
- The fixed test includes both useful paraphrase matches and failures; a top-ranked result is not proof that an item has the requested ability.
- Preserve model revision, dtype, pooling, normalization, input text, and batching when reproducing rankings.
Use embeddings to find candidates, not to decide what an item can do
A player may ask for something that helps them cross a chasm without remembering the name “climbing rope.” Sentence embeddings offer a way to rank item descriptions by meaning. They do not establish whether an item is usable, owned, unlocked, or capable of satisfying a game rule. Keep those decisions in explicit data and code.[3]
This guide tests that boundary with eight invented inventory items and seven fixed English queries. The useful result is not a claim that semantic search always works: the actual run contains a near tie, an unavailable ability, and a negation failure. All inputs, vectors, rankings, and reproduction code are linked below.

The Hugging Face resource and the exact setup
Xenova/all-MiniLM-L6-v2 packages ONNX weights for Transformers.js. Its card demonstrates mean pooling and normalization for sentence embeddings. The original Sentence Transformers card describes a 384-dimensional representation and lists semantic search among its intended uses. Both cards identify Apache-2.0 licensing. This is an established resource, checked on October 3, 2026, not a newly announced model.[1][2]
The experiment used Transformers.js 3.8.1, ONNX Runtime Node 1.21.0, q8 weights, and the CPU on Linux x64 under Node 24.19.0. The model revision was 751bff37182d3f1213fa05d7196b954e230abad9. The downloaded model_quantized.onnx file was 22,972,370 bytes; that is one file’s size, not total download size or memory usage. The report records all four downloaded files and their SHA-256 hashes.
The eight descriptions and seven queries were encoded together in that order. Mean pooling and L2 normalization yielded fifteen 384-value vectors. The ranking script computes their dot products; for normalized vectors this corresponds to cosine similarity. The semantic-search documentation describes this normalized-vector approach.[3]
This execution used the Node binding, not a browser. Microsoft documents the Node package and its platform backends separately. Transformers.js also documents browser execution through WASM or WebGPU, but those paths were not exercised here. Do not reuse these results as browser timing or compatibility evidence.[5][4]
What the fixed queries actually returned
The descriptions cover a lantern, rope, healing potion, key, shield, frost wand, apple, and fishing rod. Only each item’s descriptive sentence was embedded: IDs and display names remained separate metadata. The fixture was written before inference and was not rewritten to improve the outcome. Two consecutive initial runs produced byte-identical rankings and vectors.
| Query | First candidate | Score | Interpretation |
|---|---|---|---|
| something to help me see in a cave | Iron key | 0.321754 | Lantern was second, almost tied |
| recover after being injured | Healing potion | 0.342371 | Matches the declared recovery use |
| get across a deep chasm | Climbing rope | 0.317745 | Matches the declared crossing use |
| open a locked door | Iron key | 0.473133 | Matches the declared unlocking use |
| breathe underwater | Healing potion | 0.131016 | No item grants underwater breathing |
| LAN-001 | Wooden shield | 0.155848 | Exact ID was absent from embedded text |
| a tool that is not a light source | Brass lantern | 0.506932 | Top result contradicts the negation |
For the cave query, the key scored 0.321754 and the lantern 0.321741, rounded to six decimals. Their unrounded difference was only about 0.000013. A UI that displays three decimal places would show 0.322 for both while still choosing one first. The ranking is reproducible in this setup, but that tiny margin is a poor basis for automatically choosing an item.
The injury, chasm, and locked-door queries each put the intended item first. That demonstrates useful matches in this fixture; it does not establish an accuracy rate for real player language. The seven prompts were deliberately selected, the catalog is tiny, and no representative evaluation set or keyword-search baseline was measured.
Three boundaries a game should enforce explicitly
First, nearest-neighbor search always has a nearest candidate when the catalog is nonempty. “Breathe underwater” returned the healing potion even though its description only promises recovery. A ranker cannot add a missing capability. Use a separate capability check before exposing an actionable result, and allow a clear “no suitable item” state.
Second, an inventory code is a key, not a semantic question. LAN-001 ranked the wooden shield first in this description-only experiment. Route a recognized exact ID to a deterministic lookup before semantic retrieval. The failure does not show that the model forgot an ID: that ID was never part of the embedded item text.
Third, a negative phrase is not a reliable exclusion filter. The request for a tool that is not a light source ranked the lantern first at 0.506932. That score was higher than several useful matches above. A single score threshold would not solve this example. If the player selects “exclude light sources,” apply that constraint to a structured field instead of hoping a vector encodes the rule.
A practical design for an inventory or asset-browser prototype
An implementation worth trying is: exact ID lookup, explicit filters, semantic candidate ranking, then a visible shortlist with descriptions. Keep inventory ownership, quest state, allowed equipment slots, and permissions outside the embedding score. This is a proposed design, not a game integration tested for this article.
For a 3D asset browser, descriptions can index design intent such as “a worn brass lamp for a cave scene.” Store the actual asset ID, file location, rights record, and validated technical properties separately. Retrieval can help someone find a prop; it cannot certify that the mesh is rigged, its license fits the project, or it meets a runtime budget.
Before shipping, collect real permitted queries and label acceptable results, ambiguous requests, and no-match cases. Compare against the exact-name and tag search you already have. Keep failures in the evaluation set. Test different descriptions, model choices, and precision settings as new experiments rather than silently replacing the published fixture.
Reproduce the result before adapting it
Save catalog.json and run-search.mjs from the evidence links into the same directory. In a separate Node project, install @huggingface/transformers@3.8.1, then run node run-search.mjs. On Linux CPU-only setups, ONNX Runtime supports skipping its optional CUDA download with ONNXRUNTIME_NODE_INSTALL_CUDA=skip during installation. The script fetches the pinned public files, checks the fixture hash, writes embeddings.json and search-results.json, and disposes the pipeline. No API key or remote inference service is used.[6]
To inspect the arithmetic without loading a model, use the published vectors and recompute the dot product between a query vector and each of the first eight vectors. Preserve the full scores before rounding. Re-running inference with different inputs, batching, runtime, or precision can change the numbers; record those differences instead of treating a new run as identical.
The original model card says inputs longer than 256 word pieces are truncated by default and describes an English model. This short-English fixture does not establish behavior for long lore entries, Korean queries, multilingual search, or arbitrary negation. Evaluate those separately. The downloadable fixture contains original fictional descriptions; the package redistributes neither model weights nor third-party game assets.[2]
Evidence used
- Fixed fictional catalog and seven queries
Original eight-item fixture, including intended matches and deliberately ambiguous or unavailable requests.
- Complete CPU retrieval results
All rankings, actual scores, model and runtime versions, input hash, and downloaded model-file hashes.
- Actual pooled embeddings
The fifteen 384-dimensional vectors from the fixed run, permitting independent score recomputation.
- Reproduce the local experiment
Original Node.js CPU script; install Transformers.js 3.8.1 separately. Downloads the pinned public model, never calls an inference service.
- Reproduction instructions and code license
Installation steps, explicit experiment limits, MIT license for the original script, and fictional-fixture reuse terms.
Sources
- 1.Pinned MiniLM ONNX model card — Hugging Face. Accessed 2026-10-03.
- 2.MiniLM model card: intended uses, length limit, and license — Sentence Transformers. Accessed 2026-10-03.
- 3.Sentence Transformers semantic-search documentation — Sentence Transformers. Accessed 2026-10-03.
- 4.Transformers.js 3.8.1 documentation — Hugging Face. Accessed 2026-10-03.
- 5.ONNX Runtime Node.js binding and supported backends — Microsoft. Accessed 2026-10-03.
- 6.ONNX Runtime 1.21.0 installation script: optional CUDA download — Microsoft. Accessed 2026-10-03.
Keep asset creation and asset validation connected
Explore Goblin3D’s public features, then maintain clear descriptions and verification records for the assets you use.
Explore Goblin3D featuresRelated Goblin3D guides and tools
Sources, product facts, and original evidence were checked before publication.