convex-testing-interface
Safe HaskellSafe-Inferred
LanguageHaskell2010

Convex.ThreatModel.LargeData

Description

Threat model for detecting Large Data Attack vulnerabilities.

A Large Data Attack exploits permissive FromData parsers in Plutus validators that ignore extra members when deserializing Constr, List or Map data. If a validator's datum parser only reads the members it expects and ignores additional ones, an attacker can "bloat" the datum with extra members while preserving the validator's interpretation.

Consequences ==

  1. Increased execution costs: Processing bloated datums wastes CPU/memory execution units, making transactions more expensive.
  2. Permanent fund locking: If the datum is bloated sufficiently:
  • Deserializing the datum may exceed execution unit limits
  • The transaction required to spend the UTxO may exceed protocol size limits

In these cases, the UTxO becomes permanently unspendable and funds are locked forever with no possibility of recovery.

Root Cause ==

unstableMakeIsData and makeIsDataIndexed generate parsers that use wildcard patterns for constructor fields:

case (index, args) of
  (0, _) -> MyConstructor  -- The "_" ignores ALL extra fields!

This means Constr 0 [] and Constr 0 [junk1, junk2, ..., junk10000] both parse to the same value, allowing attackers to inject arbitrary data.

The same hole exists in the other two container shapes. A list-encoded datum is a List rather than a Constr on-chain: a parser that takes the elements it expects off the front of the list (pasList plus positional access) ignores every element after them, so List [a, b] and List [a, b, junk1, ..., junk10000] also parse to the same value. A Map datum is read by key, and PlutusTx.AssocMap.lookup returns the first matching entry, so appending entries under unused keys leaves every lookup answering exactly as it did before.

Mitigation ==

A secure validator should either:

  • Use strict manual FromData instances that check field count exactly
  • Validate the datum hash matches an expected value
  • Check datum structure explicitly in the validator logic

This threat model tests if a script output with an inline datum still validates when additional members are appended to the datum's container structure (see bloatData). If it does, the validator has a Large Data Attack vulnerability. A datum that is a bare atom has nothing to append to, so the attack is skipped as a failed precondition rather than asserting against an unmodified transaction.

Synopsis

Documentation

largeDataAttack :: ThreatModel () Source #

Default large-data attack. The number of injected fields is drawn per transaction from a curated range, so QuickCheck explores the parameter space and shrinks counterexamples toward the smallest triggering value.

largeDataAttackWith :: Int -> ThreatModel () Source #

Large-data attack with a fixed field count. Keep using this for deterministic regression tests and golden seeds.

largeDataAttackWithGen :: Gen Int -> ThreatModel () Source #

Large-data attack parameterised by a generator for the number of extra members injected into the target inline datum. This is the primitive the other two forms delegate to.

bloatData :: Int -> ScriptData -> ScriptData Source #

Bloat a ScriptData value by appending n extra members to it.

Every container shape is handled, because each one is read on-chain by a parser that can ignore trailing members:

  • ScriptDataConstructor idx fields - what unstableMakeIsData and makeIsDataIndexed produce, parsed via Constr pattern matching. The junk fields are ScriptDataNumber 42: a generated parser either matches a fixed-length prefix and ignores the rest, or matches the exact field list and fails, so the junk's type cannot change the verdict (and a record's fields are heterogeneous, so there is no "matching" type to mirror).
  • ScriptDataList xs - what a homogeneous list datum produces, and what a list-encoded record produces (e.g. plutus-tx's makeIsDataAsList, or an Aiken type declared as a list), parsed via pasList plus element access. A parser that reads the first k elements ignores every element after them.
  • ScriptDataMap kvs - parsed via pasMap plus key lookup. Appending entries whose keys do not already occur preserves every existing lookup, because PlutusTx.AssocMap.lookup returns the first match and neither FromData nor UnsafeFromData for Map validates key uniqueness, ordering, or size.

For a list or a map the junk is derived from what is already there - a member of the list, or an existing value under a fresh key of the same shape as the existing keys (see junkMemberLike and junkEntriesFor). That matters because the two parser flavours walk a different distance: the permissive UnsafeFromData instance builds its list lazily, so junk past the member being read is never even forced, but the strict FromData instance traverses every member and yields Nothing if one fails to parse - which would make the validator reject the datum outright and have the attack report a secure contract for the wrong reason. Junk shaped like a member already in the datum parses by construction.

The mirroring only pays off where the container is homogeneous - a [ByteString], a Map PubKeyHash Integer - which is where the strict instance is FromData [a] or FromData (Map k v) and does traverse everything. A list encoding a heterogeneous record (makeIsDataAsList) is read positionally like a Constr instead: a fixed-length pattern rejects any appended member whatever its type, and a prefix pattern never forces one, so no choice of junk changes that verdict. Mirroring is never worse than a constant, so both cases take the same path. (An empty list or map gives nothing to mirror, so the junk falls back to ScriptDataNumber 42 and a strict parser expecting some other member type will reject it.)

The atoms ScriptDataNumber and ScriptDataBytes are returned unchanged - they have no members to append to. largeDataAttackWithGen fails its precondition on an unchanged result, so those shapes are reported as skipped rather than silently asserted against an untouched transaction.