Hallucinated package names are a supply-chain risk: a model invents a plausible import, an attacker registers that name, the next developer or agent installs it. Published measurements put the hallucination rate between 4.6 and 6.1 percent on 2026 frontier models for Python and JavaScript. One set of 127 names was invented identically by all five models tested [1]. Every study so far covers PyPI and npm.
We re-read the 500 outputs of our pre-registered study [2]. That study covered 50 backend tasks in money, time, idempotency and access, five models, each task generated once plain and once with a one-page specification. We asked two questions. Did any output import a package that does not exist? And did the specification change what the models imported?
Imports were extracted from all 500 Python files with the ast module. Where a file did not parse as pure Python (389 files carry surrounding prose), a line-level pattern was used instead. Import names were mapped to distribution names where they differ: confluent_kafka to confluent-kafka, flask_sqlalchemy to Flask-SQLAlchemy, kafka to kafka-python, flask_login to Flask-Login. Each distribution was then checked against the PyPI JSON API on 29 September 2026, and again on 6 October 2026 with the same result. The script and the two result files are linked at the end of this note.
Seventeen distinct third-party import names appear across the 500 outputs: flask, fastapi, uvicorn, werkzeug, pydantic, requests, pika, confluent_kafka, flask_sqlalchemy, starlette, kafka, stripe, flask_login, bcrypt, boto3, botocore and psycopg. All seventeen exist on PyPI. Hallucinated package names: zero.
The second question produced the finding. Third-party imports appear in 65 of the 250 plain outputs and in 0 of the 250 specification outputs. The split holds for every model:
| Model | Plain outputs with third-party imports | Specification outputs |
|---|---|---|
claude-sonnet-5 | 16 of 50 | 0 of 50 |
deepseek-v4-pro | 14 of 50 | 0 of 50 |
gemini-3.1-pro-preview | 13 of 50 | 0 of 50 |
qwen3.8-27b | 9 of 50 | 0 of 50 |
gpt-5.6-sol | 13 of 50 | 0 of 50 |
The specification contains one relevant line: “Stack: Python 3.12, standard library only unless the task says otherwise.” No task said otherwise. The line held in 250 of 250 outputs.
It shows that a stated dependency constraint, written as a requirement, was followed by five models at the first turn without exception. It removed the entire third-party dependency surface from the generated code. Where there is no third-party import there is nothing to slopsquat.
It does not show that these models hallucinate less than others. Our tasks were designed to be solvable with the standard library. The plain outputs therefore reached for common, well-known packages (a web framework, an HTTP client, a queue client) rather than obscure ones. Zero hallucinations in 17 names is consistent with the published rates given the sample, not evidence against them. It also does not show what happens in Java or C#, where dependency coordinates are richer and where no measurement exists.
A pre-registered study of dependency hallucination in Java (Maven Central) and C# (NuGet), the two ecosystems that run most regulated back-office systems and the two nobody has measured. The design will use dependency-heavy tasks, the same five vendors, and a resolver that checks every declared coordinate against the registries. The hypothesis carried over from this note: a one-line dependency constraint cuts the surface to zero there too.
Everything needed to re-run the check is here. The script takes the results/raw folder of the dataset archive and prints the table above.
[1] The Range Shrinks, the Threat Remains: Re-evaluating LLM Package Hallucinations on the 2026 Frontier-Model Cohort, arXiv:2605.17062, 2026.
[2] S. Dhuri, Specification Before Generation: A Pre-Registered, Five-Model Paired Evaluation of a Specification Frame for LLM-Generated Code in Money, Time, Idempotency, and Access Tasks, arXiv:2609.23270, 2026. Dataset: doi:10.5281/zenodo.22598204.