TryHackMe: AI Models & Data
A complete walkthrough of the TryHackMe room AI Models & Data, exploring training data provenance, model building risks, the fine-tuning inheritance problem, model cards, and HuggingFace supply chain audits.
Overview
AI Models & Data is the third room in TryHackMe’s AI Security path, following The Building Blocks of AI and AI Security Threats. While previous rooms examined model architectures and runtime attacks like prompt injection, this room focuses on the upstream supply chain.
Before a model serves an inference request, decisions regarding dataset collection, provenance, pruning, quantisation, and fine-tuning shape its security posture. This walkthrough covers data supply chain integrity, how sensitive data becomes permanently embedded into model weights, the risks inherited during fine-tuning, and how to conduct a model security audit on open repositories like HuggingFace.
Task 1: Introduction
Every AI model is fundamentally a reflection of its training dataset. Security risks do not begin at deployment time; they originate in data collection pipelines that are often undocumented, untracked, and unaudited.
Learning Objectives
- Understand data provenance and supply chain vulnerabilities in AI training pipelines.
- Identify how Personally Identifiable Information (PII) and live credentials become baked into model parameters via web scraping.
- Evaluate the security trade-offs introduced by training decisions such as overfitting, post-training quantisation, and federated learning.
- Analyze the inheritance problem where fine-tuned models carry upstream vulnerabilities and eroded alignment from base models.
- Audit third-party models using Model Cards and evaluate supply chain risks on platforms like HuggingFace.
Task 1 Questions and Answers
- Question: I understand the learning objectives and am ready to learn about AI models and data!
Answer:No answer needed
Task 2: Where Does the Data Come From?
Training modern foundation models requires massive text corpora. GPT-3 required approximately 570GB of filtered text, while modern open-weight models like DeepSeek-V3 and LLaMA 4 train on 14 to 40 trillion tokens. To reach this scale, engineering teams rely on broad data collection categories rather than manual curation.
The Four Data Sourcing Buckets
| Source Category | Description | Trust Profile & Risk |
|---|---|---|
| Web Scraping | Automated crawls of open web pages, forums, public repositories, and social platforms. | Low. No curation, unversioned, high rate of toxic content, spam, and scraped secrets. |
| Licensed Datasets | Data acquired through commercial agreements (e.g., Reddit or news archives). | Medium. Usage terms are often ambiguous, and end users rarely consented to AI training. |
| Synthetic Data | Artificial text or code generated by existing frontier models to train new models. | Variable. Roughly 12% of fine-tuning datasets now incorporate synthetic text, risking model collapse. |
| Internal Corpora | Proprietary knowledge bases, CRM tickets, codebases, or internal documentation. | Higher. The organization controls access, but faces liability if sensitive records are exposed. |
The Dominance of Common Crawl
The primary foundation of modern LLM training is Common Crawl, an open web archive. GPT-3 sourced roughly 60% of its pre-training tokens from filtered Common Crawl snapshots. DeepSeek and LLaMA model families rely heavily on it as well.
The security exposure lies in the filtering process. Automated heuristic filters frequently miss embedded secrets, structured databases, and confidential user conversations.
1
2
3
4
[ Open Web / Internet ] ──> [ Common Crawl Scraper ] ──> [ Imperfect Heuristic Filters ]
│
▼
[ Baked into Model Weights ] <── [ Unaudited Training Loop ] <── [ Unfiltered Secrets & PII ]
The Problem of Data Provenance
Data provenance answers three fundamental questions for any training artifact:
- Where did the data originate?
- When was it collected?
- Has it been modified or tampered with since collection?
In modern AI pipelines, provenance is rarely maintained. The Data Provenance Initiative audited over 1,800 popular AI datasets and found:
- Over 70% of dataset licenses on hosting platforms were marked as “Unspecified”.
- Of the datasets that had license labels, 66% were miscategorized, typically listed as more permissive than the actual legal terms.
Software Bills of Materials vs ML-BOM
In traditional application security, Software Bills of Materials (SBOMs) became standard after supply chain incidents like SolarWinds. The machine learning counterpart is the ML-BOM (Machine Learning Bill of Materials). An ML-BOM documents:
- Source URLs and dataset lineage.
- Data collection timestamps and filtering criteria.
- Licensing restrictions and intellectual property boundaries.
- Categorization of scrubbed PII and known data gaps.
PII and Credentials in the Pipeline
Once sensitive information is processed during pre-training, it becomes mathematically encoded into the floating-point weights of the neural network. There is no simple command to delete a single fact or credential from a trained model without retraining or complex weight editing.
In December 2024, Truffle Security analyzed a 400TB Common Crawl snapshot containing 2.67 billion web pages. Their audit uncovered nearly 12,000 live, verified API keys, database credentials, and private passwords. When models train on such data, attackers can use targeted prompt extraction techniques to extract those credentials verbatim.
Task 2 Questions and Answers
- Question: What term describes the ability to answer where data came from, when it was collected, and whether it has been modified?
Answer:Data provenance - Question: What is the name of the most widely used public corpus that underpins essentially every major model family?
Answer:Common Crawl - Question: What is the AI equivalent of a Software Bill of Materials (SBOM), used to document dataset sources, licenses, and filtering decisions?
Answer:ML-BOM
Task 3: Building the Model: Key Concepts
The technical decisions made during model training directly influence its attack surface.
Epochs and Overfitting
An epoch represents one full pass of the training algorithm through the entire training dataset. Models typically iterate across multiple epochs to converge on optimal weights.
However, training for excessive epochs leads to overfitting. Instead of generalizing language syntax and semantic relationships, the model begins memorizing specific training sequences. From a security standpoint, overfit models are significantly more prone to training data extraction attacks, reproducing exact credentials, private names, or proprietary code snippets when prompted with partial matches.
Model Validation
To monitor generalization, engineers reserve a portion of the dataset as a validation set. This data is never shown to the model during weight updates.
- If training loss decreases while validation loss also decreases, the model is learning generalizable features.
- If training loss continues to drop while validation loss plateaus or increases, the model is actively overfitting.
Validation serves as the primary quality gate in the ML lifecycle. Deploying a model without rigorous validation guarantees unquantified production behavior.
Post-Training Optimisation: Pruning and Quantisation
Production models are frequently compressed to minimize memory consumption and inference latency on commodity hardware:
| Compression Method | Mechanism | Security Consideration |
|---|---|---|
| Pruning | Identifies and strips out redundant or low-weight parameters. | Alters output distributions post-training; rarely audited for safety deviations. |
| Quantisation | Lowers the numerical precision of model weights (e.g., converting 32-bit floats to 8-bit or 4-bit integers). | Degrades safety alignment boundaries. Backdoor defenses tested on 32-bit models often fail on quantised variants. |
Quantisation is commonly performed by third parties packaging open weights for consumer runtimes (such as GGUF or AWQ formats). Organizations downloading compressed models inherit modified safety boundaries that differ from the base model’s documented evaluations.
Federated Learning
In centralized training, raw data is uploaded to a unified compute cluster. Federated learning decentralizes this architecture: the model trains locally across distributed client devices (such as mobile phones or regional hospital nodes). Each device computes gradients locally on private data and only transmits weight updates back to a central aggregation server.
While federated learning enhances data privacy by keeping raw records on-device, it introduces an integrity challenge:
- Gradient / Model Poisoning: Adversarial participants can intentionally manipulate their local gradient updates to poison the global model, inserting backdoors or targeted misclassifications.
- Verification Difficulty: Because the central server cannot inspect the raw private training data, detecting poisoned updates among thousands of legitimate client nodes is technically difficult.
1
2
3
4
5
6
7
[ Central Model Server ]
▲ ▲
Weights│ │Weights
Updates│ │Updates (Poisoned?)
▼ ▼
[ Hospital A ] [ Malicious Node B ]
(Private Data) (Crafted Local Gradients)
Task 3 Questions and Answers
- Question: What term describes one complete pass of the training algorithm through the entire dataset?
Answer:Epoch - Question: What problem occurs when a model memorises training data rather than learning general patterns?
Answer:Overfitting - Question: What post-training optimisation technique reduces the numerical precision of model weights to cut memory and compute requirements?
Answer:Quantisation - Question: What training approach trains a model across decentralised devices, sending only weight updates rather than raw data to a central server?
Answer:Federated learning
Task 4: Pre-Trained Models & Fine-Tuning
Training frontier foundation models from scratch requires tens of millions of dollars in compute, specialized clusters, and months of engineering time. Most organizations instead adopt a pre-trained base model and apply fine-tuning.
- Pre-Trained Model: A large foundation model trained on web-scale datasets to learn general linguistic structure, reasoning, and world context (e.g., LLaMA, Mistral, GPT base).
- Fine-Tuning: The process of taking a pre-trained model and performing additional training passes using a smaller, specialized dataset (such as medical terminology, internal API schemas, or customer support dialogues).
The Inheritance Problem
Fine-tuning alters task-specific performance and stylistic tone, but it does not erase the underlying base weights. When deploying a fine-tuned model, an organization inherits all biases, security weaknesses, and latent training artifacts present in the original foundation model.
This manifests in three distinct ways:
- Safety Alignment Erosion: Research from Stanford and Princeton demonstrated that safety guardrails in aligned models can be broken by fine-tuning on as few as 10 maliciously crafted samples (costing under $0.20 via public APIs). Even completely benign fine-tuning on legitimate business text shifts probability distributions enough to gradually erode safety boundaries.
- Expanded Attack Surface Through Specialization: Cisco security research revealed that fine-tuning models on domain-specific corpora makes them measurably more susceptible to prompt injection. Models fine-tuned on financial or medical records become more compliant when an adversary frames an injection within that specific business context.
- Checkpoint Version Drift: Downstream teams rarely track the exact commit hash or checkpoint of the base model they fine-tuned against. If the upstream base model is later discovered to contain poisoned data or a latent backdoor, identifying exposed derivative models across enterprise systems becomes difficult.
Task 4 Questions and Answers
- Question: What is the process of taking a pre-trained model and continuing to train it on a smaller, task-specific dataset?
Answer:Fine-tuning - Question: What term describes a model that has already been trained on a large general-purpose dataset?
Answer:Pre-trained model
Task 5: The Black Box Problem
Traditional software binaries can be disassembled, analyzed with debuggers, and decompiled into readable code. Deep learning models cannot.
A model’s weights consist of billions of floating-point numbers. There is no line of code to inspect to understand why a model produced a specific token or classification. Red-team testing and input probing sample model behavior, but sampling cannot prove the absence of latent backdoors or unexpected trigger sequences.
Model Cards as Transparency Artifacts
Proposed by Google researchers in 2019, a Model Card serves as a standardized documentation artifact accompanying a model, acting like a nutritional label for AI systems.
A comprehensive Model Card outlines:
- Model Lineage & Details: Developer, version, architecture, and date of release.
- Intended Use: Documented operational scopes and explicitly out-of-scope use cases.
- Training Data: Corpora sources, filtering methodology, and known demographic or language gaps.
- Evaluation Benchmarks: Quantitative performance metrics across diverse evaluation sets.
- Known Limitations & Biases: Documented edge-case failures, behavioral degradation scenarios, and known biases.
- Licensing: Distribution terms and downstream commercial restrictions.
Industry Gaps in Practice
Unlike physical consumer goods, Model Cards remain voluntary. Organizations frequently omit critical sections:
- Training data sources are often marked proprietary or described in vague terms.
- Adverse evaluation results or safety trade-offs are omitted to avoid discouraging enterprise adoption.
- Checkpoints hosted on public repositories frequently lack Model Cards entirely.
In security operations, an absent or empty Model Card is an immediate supply chain red flag.
Task 5 Questions and Answers
- Question: What documentation artifact accompanies a model to describe what it is, how it was built, and where it falls short?
Answer:Model card - Question: What are the billions of floating-point numbers that make up a trained model collectively referred to as?
Answer:Weights
Task 6: Practical Model Audit
Task 6 provides an interactive lab simulating a public model repository platform modeled after HuggingFace. Anyone can register an account and publish model checkpoints on open platforms, introducing supply chain and arbitrary code execution risks.
1
2
3
4
5
6
[ Public Model Hub / HuggingFace Simulation ]
│
┌────────────────────────┼────────────────────────┐
▼ ▼ ▼
[ Unverified Uploader ] [ Legacy Pickle Format ] [ Missing Model Card ]
Suspicious domain/handle (.bin / .pt arbitrary RCE) No dataset provenance
Key Red Flags in Third-Party Model Audits
- Unverified Uploader Identity: Community uploads from unverified publishers or newly created accounts with no linked corporate domain or code repository.
- Unsafe Serialization Formats: Model weights distributed as Python pickle files (
.bin,.pt,.pkl). Pickle is inherently unsafe because deserializing an untrusted pickle archive can execute arbitrary system commands on the host machine. Production environments should mandate safe tensor formats (.safetensors). - Empty or Template Model Cards: Repositories missing training data provenance, evaluation metrics, and licensing documentation.
- Sudden Checkpoint Swaps: File listings showing recent weight replacements without accompanying changelogs or commit explanations.
Submitting the audit checklist for the suspicious repository completes the exercise and surfaces the flag.
Task 6 Questions, Answers, and Flag
- Question: Complete the exercise to get the flag!
Answer:THM{A_m0del_Stud3nt}
Task 7: Conclusion
Data integrity and model provenance form the bedrock of AI security:
- Models permanently encode vulnerabilities and sensitive data present in their training corpora.
- Post-training compression (quantisation) and decentralization (federated learning) introduce subtle integrity risks and safety degradation.
- Fine-tuning inherits all upstream vulnerabilities from base models, and benign specialization can weaken pre-existing safety guardrails.
- Open model repositories require rigorous supply chain inspection, rejecting unsafe serialization formats and unverified publishers.
Task 7 Questions and Answers
- Question: I have completed the room!
Answer:No answer needed
Key Takeaways
| Security Domain | Traditional Software Supply Chain | AI Model & Data Supply Chain |
|---|---|---|
| Inventory Tracking | Software Bill of Materials (SBOM) listing package dependencies | Machine Learning Bill of Materials (ML-BOM) detailing datasets and filtering |
| Integrity Checks | Package hashes and cryptographic code signing | Model weight hashes, dataset provenance checks, and safetensors validation |
| Vulnerability Remediation | Patching a vulnerable dependency or updating code lines | Retraining, fine-tuning corrective layers, or implementing inference guardrails |
| Documentation Standards | README, CHANGELOG, and OpenAPI specifications | Standardized Model Cards with explicit limitations and bias evaluations |
| File Format Security | Sandboxing executable binaries and checking shared libraries | Avoiding arbitrary Python pickle archives (.bin) in favor of .safetensors |
Summary of Questions, Answers, and Flags
| Task | Question | Answer |
|---|---|---|
| Task 1 | I understand the learning objectives and am ready to learn about AI models and data! | No answer needed |
| Task 2 | What term describes the ability to answer where data came from, when it was collected, and whether it has been modified? | Data provenance |
| Task 2 | What is the name of the most widely used public corpus that underpins essentially every major model family? | Common Crawl |
| Task 2 | What is the AI equivalent of a Software Bill of Materials (SBOM), used to document dataset sources, licenses, and filtering decisions? | ML-BOM |
| Task 3 | What term describes one complete pass of the training algorithm through the entire dataset? | Epoch |
| Task 3 | What problem occurs when a model memorises training data rather than learning general patterns? | Overfitting |
| Task 3 | What post-training optimisation technique reduces the numerical precision of model weights to cut memory and compute requirements? | Quantisation |
| Task 3 | What training approach trains a model across decentralised devices, sending only weight updates rather than raw data to a central server? | Federated learning |
| Task 4 | What is the process of taking a pre-trained model and continuing to train it on a smaller, task-specific dataset? | Fine-tuning |
| Task 4 | What term describes a model that has already been trained on a large general-purpose dataset? | Pre-trained model |
| Task 5 | What documentation artifact accompanies a model to describe what it is, how it was built, and where it falls short? | Model card |
| Task 5 | What are the billions of floating-point numbers that make up a trained model collectively referred to as? | Weights |
| Task 6 | Complete the exercise to get the flag! | THM{A_m0del_Stud3nt} |
| Task 7 | I have completed the room! | No answer needed |
