Sécurité de la chaîne d'approvisionnement LLM : Protéger l'IA d'entreprise
The discipline is still maturing, but the threat surface is not waiting. Start with model provenance verification and work outward through every layer of your stack.
Inlock focus
Inlock AI applique la provenance vérifiée des modèles, des environnements d'inférence isolés et des pistes d'audit complètes pour que les entreprises puissent faire confiance à chaque couche de leur pile LLM.
LLM Supply Chain Security: Protecting Enterprise AI From Model to Inference
The software supply chain attack surface has expanded dramatically with the adoption of large language models. Enterprises that deploy LLMs are not just running code—they are importing billions of parameters trained on unknown datasets, served through third-party APIs, orchestrated by open-source middleware, and embedded into critical business workflows. Each of these layers introduces a distinct class of risk that traditional application security frameworks were never designed to handle.
This is the LLM supply chain problem, and it remains one of the least-discussed attack surfaces in enterprise AI security.
What Is the LLM Supply Chain?
In classical software engineering, supply chain security focuses on the integrity of open-source libraries, build pipelines, and deployment artifacts. The SolarWinds and Log4Shell incidents crystallized how a single compromised dependency can cascade into thousands of breached organizations.
LLMs introduce an analogous but structurally different set of dependencies:
- •Foundation model weights: sourced from public repositories (Hugging Face, Ollama registries), academic institutions, or commercial providers
- •Training data lineage: often opaque, aggregated from web crawls, licensed corpora, and synthetic pipelines
- •Fine-tuning and RLHF layers: applied by a second or third party, potentially introducing behavioral modifications
- •Inference runtimes: vLLM, TGI, Triton, llama.cpp—each with distinct security profiles
- •Orchestration middleware: LangChain, LlamaIndex, and similar frameworks that mediate between your application and the model
- •Plugin and tool integrations: retrieval systems, APIs, and function-calling surfaces that the LLM can invoke at runtime
- •Serving infrastructure: cloud GPU instances, container registries, reverse proxies, and load balancers
A breach or compromise at any single layer can result in data exfiltration, model manipulation, IP theft, or the silent poisoning of outputs that influence consequential business decisions.
Layer 1: Model Provenance and Weight Integrity
The most foundational risk in LLM deployment is the integrity of the model weights themselves. Unlike a compiled binary that can be hash-verified against a release checksum, LLM weights are opaque parameter matrices. You cannot easily inspect them for backdoors, biases deliberately introduced during training, or data poisoning artifacts.
Model poisoning is a well-documented attack vector in academic literature. An adversary with access to a training pipeline can embed trigger phrases that cause the model to behave maliciously when those phrases appear in prompts—while behaving normally otherwise. This is sometimes called a trojan model or backdoor attack.
Enterprises must treat model weights with the same rigor they apply to third-party code:
- •Verify checksums and signatures for every model artifact before deployment. Repositories like Hugging Face now support model cards and SHA256 hashes, but signature verification via sigstore or similar tooling is not yet standard practice.
- •Prefer models with published training data transparency reports. Models trained on fully documented, auditable corpora present a substantially smaller risk surface.
- •Run behavioral red-teaming on newly adopted models before production deployment—specifically probing for trigger-phrase responses, jailbreak resistance, and unexpected capability edges.
- •Implement model versioning and rollback policies so that if a compromised or misbehaving model version is detected, you can revert to a known-good artifact within a defined RTO.
Layer 2: Fine-Tuning and RLHF Supply Chain Risks
Many enterprises adopt foundation models and then apply fine-tuning to adapt the model to domain-specific tasks. This introduces a second supply chain node. If fine-tuning is performed by a managed service provider or an internal team with insufficient controls, the fine-tuning process itself becomes an attack surface.
Risks here include:
- •Data poisoning during fine-tuning: adversarially crafted training examples that shift model behavior in subtle, targeted ways
- •Adapter and LoRA weight tampering: lightweight fine-tuning adapters can be swapped post-training without modifying the base model, making tampering difficult to detect through weight comparison alone
- •RLHF reward model manipulation: if a third-party human feedback provider is used, malicious or careless annotators can systematically bias the reward signal
Mitigations require treating fine-tuning runs as auditable, reproducible artifacts: pinned dataset versions, cryptographic hashes of training splits, signed run manifests, and post-training behavioral regression suites.
Layer 3: Inference Runtime Vulnerabilities
Once the model is in place, it must be served. The inference runtime is another layer where security gaps emerge. Popular runtimes like vLLM, Text Generation Inference (TGI), and Ollama have all received CVEs related to authentication bypass, arbitrary code execution through crafted requests, and unauthorized model access through misconfigured API endpoints.
Key hardening practices:
- •Never expose inference runtimes directly to the internet. All access should pass through an authenticated API gateway with rate limiting and request validation.
- •Pin runtime container image digests rather than floating tags. A
latesttag on your inference container is a supply chain risk—it means a compromised upstream image can silently reach production. - •Enforce network-level isolation between inference nodes and the broader corporate network. The inference runtime should have no outbound internet access unless explicitly required by a documented use case.
- •Monitor GPU memory and process isolation if you are co-locating multiple model instances on shared GPU hardware. Cross-tenant GPU side-channel attacks, while theoretical at enterprise scale today, are an emerging area of concern.
Layer 4: Orchestration Framework Risks
LangChain, LlamaIndex, and comparable orchestration frameworks have become the de facto middleware for LLM applications. They abstract away model-specific APIs, provide retrieval integrations, and simplify agent construction. They also aggregate significant attack surface.
Orchestration frameworks have demonstrated multiple critical vulnerability classes:
- •Arbitrary code execution via crafted LLM outputs: frameworks that pass LLM-generated text to Python's
eval()or similar execution contexts can be weaponized by prompt injection attacks that cause the LLM to output malicious code strings - •SSRF vulnerabilities: agent tooling that fetches URLs at the LLM's direction can be manipulated into accessing internal network resources
- •Dependency confusion: popular frameworks have dozens of transitive dependencies, each of which is a potential supply chain node
Enterprises should maintain a private mirror of approved framework versions, run automated SBOM (Software Bill of Materials) generation on every orchestration layer, and enforce dependency pinning in CI/CD pipelines.
Layer 5: Plugin, Tool, and Retrieval Integration Security
Modern LLM deployments are rarely isolated inference calls. They are connected to retrieval-augmented generation pipelines, database connectors, email systems, calendar APIs, CRMs, and custom internal tools. Each integration is a potential pivot point.
Retrieval poisoning is a particularly insidious attack: if an adversary can influence the documents stored in your vector database—through poisoned emails, crafted web pages indexed by your scraper, or direct write access—they can embed adversarial content that is retrieved and included in LLM context windows, effectively delivering a prompt injection payload through seemingly trusted data.
Controls for this layer:
- •Source-stamp all documents entering your retrieval corpus with verified provenance metadata, and filter retrieval results by trust tier at query time.
- •Implement read-only retrieval paths wherever possible. The LLM should never have write access to the same data stores it reads from.
- •Apply content validation to retrieved chunks before injection into the context window, screening for prompt injection signatures.
- •Scope tool permissions to the principle of least privilege. An LLM agent that only needs to read from a specific database schema should not have credentials that permit broader access.
Layer 6: Serving Infrastructure and CI/CD Pipeline Security
The final layer is the infrastructure that moves model artifacts from training environments to production and keeps them running. Container registries, artifact stores, model serving platforms, and CI/CD pipelines are all potential points of compromise.
Critical controls include:
- •Signed container images using Cosign or equivalent tooling, with policy enforcement that rejects unsigned images at deployment time
- •Immutable artifact storage for model weights and quantized versions, with access logging on every retrieval event
- •Separation of duties between teams that can write model artifacts and teams that can promote them to production
- •Runtime integrity monitoring that detects unexpected modifications to model files on disk during serving
Building an LLM Supply Chain Security Program
Addressing these risks requires a program-level response, not a checklist. An effective LLM supply chain security program for enterprises includes:
1. LLM-Specific SBOM Generation Extend existing SBOM practices to cover model weights, adapters, training dataset manifests, and runtime dependencies as first-class artifacts.
2. Model Acceptance Testing Gates No model—base or fine-tuned—should reach production without passing a standardized behavioral test suite that probes for known attack patterns, policy violations, and regression against prior behavior baselines.
3. Continuous Drift Monitoring Model behavior can shift over time even without explicit updates, particularly in systems that incorporate ongoing fine-tuning or retrieval from evolving corpora. Implement output distribution monitoring to detect behavioral drift before it causes production incidents.
4. Incident Response Playbooks for Model Compromise Define and rehearse the response to a suspected model tampering event: how to identify affected workloads, how to isolate and roll back, how to communicate with affected stakeholders, and how to conduct a post-incident behavioral forensic analysis.
5. Vendor Risk Assessment for AI Components Apply the same vendor risk management discipline to AI component suppliers—model vendors, fine-tuning services, inference infrastructure providers—that regulated industries already apply to critical technology vendors.
Regulatory Context
Supply chain security for AI is increasingly a regulatory expectation rather than a best practice. The EU AI Act requires high-risk AI systems to maintain technical documentation covering training methodologies and datasets. DORA requires financial entities to manage ICT supply chain risk, which now necessarily includes AI model providers and inference infrastructure operators. NIS2 extends supply chain security obligations to critical infrastructure operators, many of whom are rapidly integrating LLMs into operational workflows.
Organizations that treat LLM supply chain security as an optional hardening exercise will find themselves out of compliance with overlapping regulatory requirements within the next 18 to 24 months.
Conclusion
The LLM supply chain spans model weights, training data, fine-tuning processes, inference runtimes, orchestration frameworks, retrieval integrations, and serving infrastructure. Securing it requires extending classical supply chain security disciplines—provenance verification, SBOM generation, artifact signing, least-privilege access—to a new class of opaque, high-dimensional software artifacts.
Enterprises that deploy LLMs without a structured supply chain security program are not just accepting technical risk. They are accepting operational, regulatory, and reputational risk that grows in proportion to how deeply those models are embedded in consequential business processes.
The discipline is still maturing, but the threat surface is not waiting. Start with model provenance verification and work outward through every layer of your stack.
Next step
Check workspace readiness
Validate connectors, RBAC, and data coverage before piloting Inlock's RAG templates and draft review flows.