Home / Data & AI threat intelligence

Data · AI · Engineering

Your pipeline is part of the attack surface.

Threat intelligence for the systems moving, transforming, training and serving your most valuable data. We map who would target them, how the trust chain can be abused, and what that means for engineering teams before it becomes an incident.

Data platforms & warehouses AI & MLOps ETL / ELT pipelines Engineering abuse paths

The problem

Security teams see assets. Attackers see trust.

A modern data estate is a chain of identities, jobs, notebooks, schedulers, object stores, APIs, models, packages and automation. The danger is not only whether one component has a CVE. It is whether an attacker can abuse the relationships between them to do something valuable.

01

Pipeline abuse

ETL and ELT workflows already move sensitive data at scale. If an adversary gains the right identity or execution point, the pipeline itself can become the exfiltration mechanism.

  • Job and DAG modification
  • Scheduler and runner compromise
  • Service account reuse
  • Trusted egress paths
02

Data & model integrity

Attackers do not need to steal everything. In some environments, changing what the business trusts can be more damaging than removing it.

  • Dataset poisoning
  • Training-data manipulation
  • Model registry tampering
  • Feature-store abuse
03

Engineering identity

Developer and machine identities often bridge source control, CI/CD, cloud, notebooks and data platforms. That makes them disproportionately valuable.

  • Tokens and API keys
  • Secrets in notebooks
  • CI/CD credentials
  • Over-privileged workload identities

What we map

The real data attack surface.

We map the systems that actually create attacker leverage: not just hosts and endpoints, but the flows between engineering, data and AI services.

That means understanding where identities cross boundaries, where code becomes execution, where one trusted platform can reach another, and where normal business processes can hide malicious activity.

Snowflake Databricks Spark Airflow dbt Kafka MLflow Notebooks Vector databases Object storage Feature stores CI/CD Model APIs Cloud IAM

Method

Threat intelligence built around how the platform works.

We start with the operational reality of your environment, then connect it to actor intent, capability and plausible abuse paths.

01

Map the estate

Understand the platforms, pipelines, trust relationships, identities, data movement and critical dependencies.

02

Identify what matters

Define the datasets, models, decisions, functions and engineering processes whose compromise would create real impact.

03

Map relevant threats

Assess actors, campaigns and techniques against your technology, sector, geography and operating model.

04

Build abuse paths

Join the evidence into realistic sequences that show how legitimate engineering capability could be turned against you.

05

Test the response

Convert the strongest scenarios into tabletop exercises and control-team questions for engineering and security.

Example abuse path

Exfiltration without “breaking” the pipeline.

The strongest scenarios are usually not cinematic. They are built from ordinary privileges, ordinary automation and an attacker who understands how engineers expect the platform to behave.

Initial access

A developer credential, CI token or notebook secret is obtained through phishing, token theft, exposed configuration or a compromised dependency.

Execution

The attacker gains the ability to change a pipeline task, scheduled job, package reference or transformation step already trusted by the environment.

Privilege through trust

The job inherits access to downstream datasets, object storage, secrets or compute because the platform assumes the workload itself is legitimate.

Collection

Sensitive records are selected or duplicated during normal processing, avoiding the need for noisy bulk discovery from an endpoint.

Exfiltration

Data leaves through an allowed integration, export task, object-store replication path, model endpoint or other expected business workflow.

Why it matters

The controls may all be “green” individually. The risk lives in the trust chain between them.

Threat mapping

From actor to engineering decision.

We do not stop at “APT X uses credential theft”. We translate threat evidence into the systems and decisions that matter to your teams.

01

Actor and campaign relevance

Who has the intent and capability to target your sector, data, geography or technology stack?

02

Technique mapping

Map known techniques to the closest meaningful engineering behaviours without forcing every data or AI abuse case into a framework category that does not fit.

03

Platform translation

Translate the threat into concrete controls and trust relationships across data warehouses, orchestration, cloud IAM, MLOps and development tooling.

04

Scenario construction

Build the attack sequence your engineers can actually reason about, test and detect.

05

Decision and action

Identify what should change, who owns it, how it can be tested, and what evidence would demonstrate the risk has reduced.

AI & MLOps

AI systems add new trust relationships, not magic.

The useful question is not whether you “have AI risk”. It is where AI and MLOps extend the attack surface you already depend on.

DATA

Training and retrieval sources

Who can alter the material that models, agents or retrieval systems trust? What happens if poisoning looks like legitimate data change?

MODEL

Registries and deployment

Model artefacts, registries and deployment pipelines can become software supply-chain problems with different terminology.

IDENTITY

Agents, tools and APIs

AI-enabled workflows often introduce privileged API access, service identities and tool execution that attackers can attempt to redirect or misuse.

We separate AI-specific risk from ordinary security failure.

Prompt abuse matters where it changes authority, data exposure or downstream execution. A leaked cloud key is still a leaked cloud key. The assessment should distinguish the two rather than relabelling every established security problem as “AI security”.

Tabletops

Exercises your data and AI teams will recognise.

Generic ransomware tabletops tell you very little about how a data platform team will detect and contain a compromised scheduler, poisoned model artefact or abused service identity.

01

Pipeline compromise

A trusted orchestration job is modified and begins copying regulated data through a permitted integration while observability remains superficially normal.

02

Data integrity attack

A source feeding executive, risk or fraud decisions is deliberately manipulated. Teams must determine when they stop trusting downstream outputs.

03

MLOps supply chain

A model, dependency or package enters the environment through the development workflow and creates unauthorised access after deployment.

Deliverables

What you actually get.

Threat landscape
Relevant actors, campaigns, motives and observed behaviours mapped to your data and AI estate.
Attack surface map
Critical platforms, identities, trust boundaries, pipelines, integrations and high-value flows.
Abuse paths
Plausible sequences showing how trusted engineering capability can be turned into attacker leverage.
Scenario pack
Engineering-ready threat scenarios suitable for design review, control testing or red/purple-team planning.
Tabletop exercise
Facilitated scenario with injects, decision points, observations, ownership and remediation actions.
Executive view
A concise explanation of material exposure, confidence, business impact and priorities.

Who it is for

Built across security and engineering.

This work is useful when security understands threats but not the data platform, or engineering understands the platform but has never modelled how an adversary would abuse it.

CTI
Threat teams that need to move beyond generic sector reporting and into platform-specific relevance.
Data
Data engineering and platform teams responsible for pipelines, storage, transformation and orchestration.
AI
AI and MLOps teams operating models, registries, retrieval systems, agents and deployment workflows.
SOC
Detection teams that need scenarios grounded in the logs, identities and workflows the platform actually produces.

Why this matters

Scanning finds weaknesses. Intelligence finds reasons.

A vulnerability scanner can tell you a package is old. It cannot tell you whether compromising that package gives an adversary access to a scheduler identity that can read a regulated dataset, invoke a model endpoint and export through an approved SaaS integration.

That is a different problem. It requires understanding intent, access, trust and consequence as one system.

The question is not “what is vulnerable?” It is “what can be abused, by whom, for what outcome, and would we recognise it?”

Contact

Map the attack path before someone else does.

If you are running modern data, AI or engineering platforms and your threat model still begins and ends with endpoints and CVEs, there is a gap worth looking at.