On avoiding YAML
YAML is a popular format for configuration files, with a graph data model, extensible types, and a supposedly human-readable syntax. Yet, almost no language ships a YAML parser in its standard library (except Ruby) which suggests nobody wants to own one. That was the first flag for me. In this post I go through the pitfalls that convinced me: a complex syntax, insecure defaults, and a dependency I don’t need.
What is YAML
YAML Ain’t Markup Language™ (YAML) describes itself as a “human-readable data serialization language designed for data exchange between languages with different data structures”. YAML is popular in “declerative automation”, for example, in infrastructure as code (Docker Compose, Kubernetes, …), and CI/CD (GitHub Actions, GitLab Runners, …).
The “human-readable” claim is based on the fact that, unlike JSON, YAML avoids using brackets. This development is in line with the evolution of popular data serialization formats: JSON surpassed XML by eliminating closing tags and replacing chevrons with braces.
YAML represents a Graph, whreas XML and JSON represents trees. It has anchors and aliases that are resolved at the parser level.
What is wrong with YAML?
It is ‘human-readable’, more expressive than other data models (graph vs tree), and widely used. So why not use it? The following are the issues that I have with YAML. They don’t necessarily impair its functionality or usability.
YAML is not Human-readable
No, I never said it was ‘human-readable’. The authors claim that it is, but I vehemently disagree. Let me give you some examples:
# NOTE: flow style is similar to vanilla JSON.
# The following examples start with flow
# style and then the 'human-readable'
# block style.
# ----------------- EXAMPLE 1 ----------------- #
## Flow style
[1,2,3]
## Block style with YOLO spacing before elements
- 1
- 2
- 3
## Block style with YOLO spacing before the block
- 1
- 2
- 3
# ----------------- EXAMPLE 2 ----------------- #
# Flow style
[[[[1, 2], [3]], [[4]]]]
## Compact block style
- - - - 1
- 2
- - 3
- - - 4
## Block style with one level per line
-
-
-
- 1
- 2
-
- 3
-
-
- 4
## Block style with YOLO mix and match
- - - - 1
- 2
-
- 3
-
-
- 4
# ----------------- EXAMPLE 3 ----------------- #
# Flow style
[1, "2 - 3"]
## Because when will you ever have an extra erronous space in your code?
- 1
- 2
- 3I would say that YAML is rather ‘human-legible’ (as XML specs put it) than ‘human-readable’. You can show the examples above to someone and they would recognize each character, but I’d bet that only a few would be able to understand what it represents. A counter example would be SQL, which was purposefull designed be similar to a natural language (English) so the following example is both ’legible’ and ‘readable’:
SELECT class_name FROM teachers
WHERE last_name = "Ms. Sequel";Parsing YAML is a pain
Have you ever wondered why none of the popular programming languages (except Ruby) have YAML parsers in their standard library? I have my own theory: it’s because YAML grammar is unnecessarily complicated.
Like any other language, you have to parse YAML and map it to a data structure in the memory.
Take JSON, feed it to a parser generator (compiler-compiler), and you have your basic parser (not comparable with a standard lib implementaion, e.g., Go encoding/json).
The same does not apply for YAML, because at every step, the parser needs information about it’s surrounding that cannot be expressed as easily as it is done for a context-free grammar.
Let’s take a look at arrays (or collections in YAML).
For JSON we have the following ABNF rule:
array = begin-array [ value *( value-separator value ) ] end-arrayWhen the parser encounter begin-array (i.e., the [ character), it starts storing the values separated by the value-separator as elements of the array until end-array (i.e., ]) is reached.
For the same type in YAML (1.2.2) we have the following parametrized BNF:
[183] l+block-sequence(n) ::=
(
s-indent(n+1+m)
c-l-block-seq-entry(n+1+m)
)+
[184] c-l-block-seq-entry(n) ::=
c-sequence-entry # '-'
[ lookahead ≠ ns-char ]
s-l+block-indented(n,BLOCK-IN)Say what? We really need to break this down:
- Unlike ABNF, the rules are parametrized. For example,
n, is the parent indent: how many spaces haven been seen up to here. For the root elementnequals-1. - There are also variables with values inferred from the context.
m, for example, is the constant (within the block) that denotes the number of spaces used to indicate indentation; it can be0, it can be100. We chosem = 0for now. s-indent(n+1+m)says this many spaces are to be expected before the entry. If we are at root level, it is expected to have-1 + 1 + 0 = 0spaces.c-l-block-seq-entry(n+1+m)defines one element in the array.c-sequence-entry(i.e.,-), followed by an space (denoted by negative look ahead of a ’non-space’ char[ lookahead ≠ ns-char ]) defines a single element.s-l+block-indented(n,BLOCK-IN)accepts two parametersnandBLOCK-INand defines the element boundary or more specifically (from the specs):
The entry node may be either completely empty, be a nested block node or use a compact in-line notation. The compact notation may be used when the entry is itself a nested block collection. In this case, both the “-” indicator and the following spaces are considered to be part of the indentation of the nested collection. Note that it is not possible to specify node properties for such a collection.
So how does the parse know when the array ends? When s-indent(n+1+m) does not match anymore.
Yes, it could’ve been a closing bracket.
BUT NO!
We want ‘human-redeability’ and a closing bracket breaks human brains as it is well known.
The YAML 1.2.2 spec defines 211 productions. JSON’s entire grammar fits into 14 rules that you can hand straight to a parser generator; something no one can do with YAML specs.
It shows in the results. The YAML test matrix has only a couple of implementations passing the full suite, and widely used parsers such as PyYAML and go-yaml accept more than a dozen documents each that the spec declares invalid. In fairness, the data is from 2022, some processors target YAML 1.1 on purpose, and the suite deliberately leans on edge cases.
YAML is Unsafe by Default
One of the key features of YAML is its extensibility: you can define your own data types through tags. Those tags are then mapped to language-specific objects, much like object-relational mapping (ORM) turns rows from a relational database into objects of an object-oriented language. With an ORM, though, the mapping is written in your code, ahead of time. A YAML tag inverts that: the document decides which type the loader should construct. That inversion is the entry point for the remote code execution flaws found in several YAML implementations.
The consequences are documented. PyYAML FullLoader allowed arbitrary code execution through the python/object/new constructor, first in CVE-2020-1747 (fixed in 5.3.1), then again in CVE-2020-14343 (fixed in 5.4), because the first fix was incomplete.
SnakeYAML default constructor allowed the same thing until CVE-2022-1471, resolved by disabling global tags by default.
In both cases the fix was to take the feature away.
XML and JSON reach the same place, but only by leaving their specifications behind.
Neither format carries type information: a <person> element and a {"type": "person"} object are just markup and data.
With XML or JSON, this type of vulnerability has to be added above the parser layer.
With YAML, it was the default, in the two most widely used implementations, because the specification puts the mechanism in the format itself.
YAML has no schema language of its own
“So what?” Well, a schema is how you state what your data is allowed to look like, and then let both producers and consumers check against it. XML has XSD, and JSON has JSON Schema. YAML has neither. In practice people validate their YAML by treating it as JSON and reaching for JSON Schema. This works in part because YAML data model is a superset of JSON, yet fails when it comes to tags, anchors, non-string keys, multi-document streams, etc. Those are exactly the YAML-only features, and they’re precisely what no schema can describe.
Neither XML nor JSON had a capable schema language from the start. XML shipped with DTDs, but they had no data types and no namespace support; XSD followed in 2001. JSON Schema only became usable around 2013, a decade after JSON was introduced. In both cases, heavy use in APIs is what made schemas clearly worthwhile: validation rules stop being reimplemented per programming language and live in one declarative document instead.
YAML never grew one, and I think that’s because YAML stayed in configuration, where the only consumer is the program reading the file and validation lives in its code.
That defence doesn’t cover my case. In Xenguard, users configure components through a web GUI, so the configuration is machine-generated and machine-consumed: an interface, not a file someone hand-edits. I want to validate it before it reaches the module and fail loudly if anything is wrong. I could do that by validating YAML with JSON Schema, but then the YAML is only a syntax wrapper around a JSON document, and all it adds is a parsing layer where the file and the validator can disagree.
YAML is YAUD
Haven’t heard of YAUD? Well, I just invented it. It stands for “Yet Another Unnecessary Dependency”.
None of the popular programming languages, except Ruby, ship a YAML parser in their standard library, so you have to opt for a third-party implementation, and hope someone keeps maintaining it.
That hope has not aged well.
The two most widely used implementations in Go and Rust were both abandoned by their original authors within about a year, and the YAML Organization had to step in and take over development and maintenance (see yaml_serde and go-yaml).
So the deal is: an extra dependency, its churn, and its attack surface. And at this point I’m still not sure what YAML gives me that JSON doesn’t.
YAML is not standarized
This might be the least problem with YAML but still worth a mention.
Both XML and JSON have standardized specs: W3C for XML and ECMA for JSON with both specs also published as ISO standards. There are also a number additional IETF RFCs for both. In contrast, the only normative document for YAML is git repo maintained by the YAML team.
Standardization introduces a streamlined process of proposing changes, reviewing them, getting feedback from the commiunity, etc., before introducing them.
Final words
YAML is a popular data format, especially for configuration files. It aims to be simple, yet flexible and extensible. In this post I’ve highlighted a number of its pitfalls and argued why it isn’t worth using instead of JSON, or even XML.
YAML looks easy at first glance, but it hides a complex syntax (two of them, in fact, flow and block style) which makes its parsers large and hard to get right, and that is where bugs and vulnerabilities come from.
Validating YAML in practice means reaching for JSON Schema, which works only because YAML data model is a superset of JSON. Everything YAML adds on top sits outside what the schema can describe. And those additions are exactly the features implementations have been turning off by default, after they turned out to be attack surface. What’s left is JSON with a harder grammar, plus a dependency your code base doesn’t need.