Abbreviated Quantity Construction¶
Ever since C++11 brought user-defined literals (UDLs) to the language, most C++ units libraries
have provided them out of the box for a wide variety of units. It’s a very popular feature —
after all, who wouldn’t rather write 65_mph than miles_per_hour(65)? Checking off the
“user-defined literals” box makes a library seem more complete, more appealing.
We’d like to reframe that box in terms of the need it’s actually meeting. Instead of “user-defined literals”, think of it as “abbreviated quantity construction”. Users want concise, readable ways to make quantity objects. UDLs are one way to do that, but not the only way. In fact, for quantities in particular, UDLs have serious flaws that competing approaches don’t suffer from.
This article will explain how UDLs generally work in a quantity context, and the downsides that turned Au’s reluctance into outright rejection. We’ll explain an alternative approach, unit symbols, that solves virtually all of UDLs’ problems. Finally, we’ll share the surprising discovery that put UDLs back on the table: by giving them non-quantity return types, they turned out to handle the quantity use cases better than any other library on the market! We’ll close with guidance on which tool to reach for, when.
UDLs for quantities¶
The basic definition of a UDL is simple. Although the spelling of the function name is somewhat
arcane (for example, the _m literal is spelled operator""_m), it is, ultimately, just a
function. As the docs explain, the set of permissible parameter types is
heavily restricted. For quantity literals, the most relevant options are long double for floating
point literals, and unsigned long long for integer literals.
Let’s define one now. We’ll write what we’d get if Au were to follow the common style in other
libraries, using meters as an example:
With this, we can write 3.5_m, and we’ll get a QuantityD<Meters> object with a value of 3.5.
This is a huge benefit for migrating callsites, especially when functions have many parameters.
Having many literals in a callsite is common in unit test files. You can migrate those tests by
just adding a little unit suffix on each number. This kind of change will be a lot easier to review
at a glance, and much less likely to overflow the line width than spelling out all the unit names in
full.
Small nuisances¶
Already, we might notice some aspects that feel a little off. For one thing, we’ve got a
static_cast to double for the return type. That’s because raw floating point literals are
double by default (decltype(3.5) is double), and we want our literals to match the original
literals they’re replacing. So, we cast. (Note that this is fine: although the parameter type is
restricted, the return type can be whatever you want.)
In fact, it gets worse, because something like 3_m doesn’t even compile! The 3 cannot match the
long double parameter type; instead, it looks for an unsigned long long overload. So if we want
this to cover all numeric literals, we need two overloads for every unit, with the second looking
something like this:
// NOLINTNEXTLINE(runtime/int)
constexpr auto operator""_m(unsigned long long v) {
// What is `T`? Let's discuss...
return meters(static_cast<T>(v));
}
The NOLINTNEXTLINE is just something you may need to add1 depending on your linter
configuration. It’s not aesthetic, but it’s livable.
The real question here is what type we should return. On the one hand, decltype(3) is int,
which argues for an int return type. On the other hand, 3_m and 3._m look very similar, and
some libraries want to avoid the . changing the type in a meaningful way — especially when
integer-backed quantities are typically far less fleshed out than floating point ones.2 Thus,
double is also a common choice for T in this context (for example, the nholthaus units library
makes this choice).
The case against quantity UDLs¶
These oddities turn out to be just the tip of the iceberg. Mateusz Pusz (lead author of the mp-units library) was the first person we know of to make a systematic case against using UDLs for quantities. Some of the key problems he identifies include:
-
UDLs compose poorly. If you have
_mand_s, you can’t make one for_mps(meters per second) without writing an entirely new overload. Actually, make that two new overloads — one for floating point, and one for integral. That means… -
UDLs are expensive. We need two copies for every unit, so the maintenance cost is high. Additionally, within Aurora, we found that the compile time cost of having many UDLs adds up to a significant amount.
-
UDLs don’t let you pick the rep3. The real code smell with the
static_cast<T>is that the UDL author is forced to chooseT. That choice should belong to the user! With UDLs, if we supportdouble, we cannot supportfloatorlong doubleat the same time. -
UDLs only work for literals. If you have a raw number in a variable, you’re out of luck: your only option is to spell out the full unit name.
As a historical point, Aurora’s internal units library (the precursor to Au) always had UDLs, the same as every other units library at the time. When it came time to open source Au, something about the feature didn’t feel right. Although we couldn’t quite put our finger on the problem, we deliberately omitted UDLs from the public library. It wasn’t until our collaborations with mp-units that we realized we had dodged a bullet. Fortunately, that wasn’t the end of it — mp-units also showed us an alternative approach that fixed all of the problems we listed above!
Unit symbols¶
A “unit symbol” is a simple idea. It’s an object that can multiply or divide with both raw numeric variables and quantity objects. The result is always a quantity object, whether or not the input already was. And the effect is just to change the units of the variable.
Here’s a brief example. If m is a unit symbol for meters, and s is a unit symbol for seconds,
then:
3.5fis a raw number.3.5f * mis a quantity of meters: we changed a raw numeric type into a quantity type.3.5f * m / sis a quantity of meters per second: we changed one quantity type (3.5f * m) into another, with different units.
Unit symbols are monovalue types. They’re empty, so the “multiplication” or “division” never has any runtime cost. Instead, it just tells the compiler how to change the units of some other variable that does have a runtime value.
Let’s see how all of our UDL problems melt away with unit symbols:
-
Unit symbols compose naturally. We’ve already seen how
m / sis a symbol for meters per second, composed on the fly. We can also apply prefixes inline:kilo(m)does just what it looks like, and we can even make it into a new named symbol locally in a file for even better readability.4 -
Unit symbols are cheap. Just one definition covers every rep, so we’d expect it to be twice as fast as the two-definition UDL approach. In practice, we’ve found it to be noticeably faster than even a single overload. Better maintenance, lower compile time cost.
-
Unit symbols work with any rep. It’s perfectly natural.
3.5 * mgives us a rep ofdouble, and3.5f * mgives us a rep offloat. In fact,Eigen::Vector3d{1.0, 2.0, 3.0} * meven gives us a rep ofEigen::Vector3d! -
Unit symbols support variables. You can write
my_legacy_raw_numeric_variable_m * m, and see at a glance (..._m * m) that we’ve gotten the units right.
It’s not as if UDLs have no advantages. 3.5_mps has a more concise, unified appearance than
3.5 * m / s. (Even if we defined an ad hoc mps symbol, to be as brief as possible, 3.5_mps is
still more concise than 3.5 * mps would be.) Additionally, the multiplicative syntax might make
new users (wrongly) fearful of runtime costs. However, on adding up all of the costs and benefits,
it’s not even a close call: unit symbols win by a country mile.5
Au’s Constants¶
This next section might seem like an odd detour. And it is, in many ways, but it’s what led us to reconsider UDLs, and discover some surprising strengths.
One of Au’s key hidden strengths is the Constant template type. We originally
introduced it to support fundamental physical constants such as the speed of light, or Planck’s
constant. Like unit symbols, a Constant is a monovalue type: it can only
ever hold one value, so we always know that value. This gives it a versatile superpower: the
perfect conversion policy.
To understand what this means, contrast Constant with Quantity. A quantity is not a monovalue
type. It wraps some ordinary numeric type, which means we can’t reason at compile time about what
the value is. All we can do is reason about what it might be. When dealing with conversion
risks, we think about how many values are lossy, and which values are lossy, and then we make
the best blanket decision we can: allow the conversion, or forbid it by default. But these are
heuristics, and they have false positives and false negatives.
Let’s take a concrete example: a CPU has a word size of 64 bits, and we want to pass that constant
to a function expecting bytes. If we reach for a Quantity, we have:
constexpr auto word_size = bits(64);
// ⚠️ Doesn't compile: truncation risk too high!
constexpr auto word_size_bytes = word_size.as(bytes);
This is annoying, but also perfectly reasonable: a quantity of bits generally won’t be exactly
representable in bytes. We can circumvent this by passing a conversion risk policy parameter,
telling the compiler (and the reader!) that this particular truncation risk is not a concern:
word_size.as(bytes, ignore(TRUNCATION_RISK)) unblocks the conversion.
While this works, it’s less than perfectly satisfying. After all, we know that this particular
value will not actually truncate, even though most integer values would. If only we could make
use of that knowledge directly! This motivates us to reach for ad hoc Constants (as opposed to
constants of nature), which means we need to know how to construct them.
Making Constant¶
The main way to make an ad hoc Constant is with the make_constant() function. Its single
parameter is a unit slot, so it supports a wide variety of inputs. Here’s a natural solution for
our current example:
constexpr auto word_size = make_constant(bits * mag<64>());
// ✅ Perfect conversion policy verifies: lossless conversion!
constexpr auto word_size_bytes = word_size.as<int>(bytes);
The one change we had to make was to add the <int> template parameter to as(). This makes
sense: Constant has no associated rep at all, so we need to tell it what type to store it in.
Once we do, Constant has everything it needs for its perfect conversion policy. If the value fits
in the target unit-and-rep, it allows it; if not, it doesn’t. If we replaced 64 with 65, it
would — correctly — fail to compile!
Au 0.6.0 put the flexibility of Constant into overdrive. Besides multiplying and dividing them,
we can add, subtract, compare, and even % them, and Au will generate brand new types on the fly
to represent the results. This means we’re going to be making Constant a lot more often in
modern Au — which means every flaw and nuisance will be magnified.
Making Constant better¶
The perfect conversion policy is amazing, but the readability took a real hit. Instead of just
bits(64), we now find ourselves writing make_constant(bits * mag<64>()). And it gets worse for
fractional numbers — a Quantity like 1.5 * s turns into make_constant(seconds * mag<3>() /
mag<2>()) if we upgrade it to a Constant. We don’t want to force users to do math in their heads
just to understand our intent!
Fortunately, there were still more powerful varieties of UDLs that we hadn’t yet considered.
Instead of simple functions taking long double or unsigned long long, we can move the literal
into a variadic template parameter pack. Note that the template parameters will be different for
every distinct sequence of characters. This means each individual literal can have a different
return type, even with the same unit suffix!
One type per number is just what we need to support a monovalue type like
Constant. Here’s what the signature looks like:
template <char... Cs>
constexpr auto operator""_s() {
// Figure out how to parse digits, exponents, decimal, separators, etc. 🤔
// Return an appropriate `Constant` type.
}
The details aren’t very enlightening, so we’ve omitted them here. The key takeaway is what you can
do with this tool: instead of make_constant(seconds * mag<3>() / mag<2>()), or even 1.5 * s,
we can now simply write 1.5_s. We get all the superpowers of the first construct (it’s a
Constant, not a Quantity), with even more conciseness than the second!
Remember: even when these values look like floating point numbers, they are not. Instead, every
Constant literal is an exact rational number. When you see 1.2e-3_s, it means exactly (12
/ 10,\!000) seconds, even though no floating point number can represent this value.
Let’s drive this point home further. Most units libraries will forbid assigning a floating point
value to an integer-backed quantity, because this is usually lossy. But with Constant literals,
we can assign 1.2e-3_s to a Quantity<Micro<Seconds>, int>, and get exactly
micro(seconds)(1'200)! Au’s new UDLs give you the freedom to specify your Constant values with
a level of readability we could hardly have imagined before.
Non-Quantity quantity UDLs¶
Now that we know the power of Constant, let’s take a fresh look at the ingredients we ended up
with. We have UDLs that flexibly and concisely construct Constant objects. And we know those
Constant objects have perfect conversion policies with any Quantity types of their same
dimension. If we put them together, we see that Constant UDLs can act like the old quantity
UDLs when we pass them to APIs expecting a Quantity!
This raises the obvious question: what about all those UDL downsides we listed earlier? Let’s revisit them.
-
Poor composition: partial improvement.
- Remember that UDLs produce a
Constant, so we have full access toConstantAPIs. For compound units, we can just multiply or divide at the end:3.5_m / sis a workable substitute for3.5_mps. And for prefixed units, we can apply the prefix to the whole constant:kilo(54.3_g)certainly looks unusual, but it’s obvious that it means the same thing that54.3_kgwould have meant, and there are key use cases where the UDLs’ flexibility outweighs the strange syntax (see the atomic units example, for instance).
- Remember that UDLs produce a
-
Expensive: partial improvement.
- We’re down to a single copy, so the maintenance cost is halved. The compile time cost per overload is likely much higher, but we have a strategy to mitigate it: every UDL gets its own header file, so every translation unit includes only the UDLs it actually needs.
-
Rep choice: solved.
- The UDLs return
Constant: a monovalue type. Ironically, by having no rep, we automatically support every rep! (At least, every rep that we can check at compile time for lossiness — this is currently limited to the arithmetic reps, but see #52 for future plans.)
- The UDLs return
-
Literals only: no change, but no longer a problem.
- UDLs are still, unsurprisingly, only for literals. That’s fine, though, since this is the
only use case we have when
Constantis involved: everyConstantmust be known directly at compile time.
- UDLs are still, unsurprisingly, only for literals. That’s fine, though, since this is the
only use case we have when
Overall guidance¶
Au provides two methods for abbreviated construction — unit symbols, and unit literals (UDLs) — with different strengths and weaknesses. Here’s how they compare, criterion by criterion. (The colors reinforce the text; they use the same colorblind-friendly scheme as our library comparison matrices.)
| Unit literals (UDLs) | Unit symbols | |
|---|---|---|
| Conciseness | 3.5_m: nothing shorter is possible |
3.5 * m: still short, but more visually spread out |
| Numeric notation | Any digits, decimal point, or exponent you like --- and the result is an exact rational number | An ordinary rep value --- the same one your program would use anyway --- with the usual rounding and overflow rules |
| Convertibility | Perfect: converts to any unit and rep that can hold this exact value, and refuses every one that can't | The usual Quantity rules: based on what the type
could hold, so some individually safe conversions are still blocked |
| Non-literal values | Not supported: UDLs are for literals only | legacy_duration_s * s works fine |
Generic Quantity APIs |
Not supported: a Constant can't know which
Quantity to convert to |
Produces a Quantity directly, so there's nothing to
deduce |
| Composability | Prefixed and compound units need help: kilo(54.3_g),
55_mi / h |
Composes on the fly: kilo(m), m / s |
| Rep support | Every rep we can check for lossiness at compile time (currently the arithmetic reps; see #52) | Any rep at all, including custom ones such as
Eigen::Vector3d |
| Availability | Opt-in per unit: needs an extra au/units/literals/...
header |
Ships with the unit's own header; nothing extra to include |
| Namespace granularity | All-or-nothing in practice: the conventional
using namespace ::au::au_literals; brings in every literal you've
included |
Name exactly the symbols you want:
using ::au::symbols::m; |
Both tools bring very short names into scope
m, s, _m, _s: these are the shortest names in your program, and the most likely to
collide. A file-scope using ::au::symbols::m; or using namespace ::au::au_literals; is
fine in an implementation file (.cc, .cpp), where it affects a single translation unit. In
a header, never do either at namespace scope: the names would leak into every translation
unit that includes it. See Namespaces and includes for the full treatment.
The literals in particular are still quite new, and our best practices could evolve as we learn more in practice. Nevertheless, here’s our best current understanding.
-
Definitely reach for literals (UDLs) when…
- …you are making a
Constant. - …you want to specify an exact value (such as a measured physical constant from CODATA) in a readable way.
- …you are making a
-
Definitely reach for unit symbols when…
- …your value is not a literal.
- …you are passing to a generic (template) API that expects a
Quantity: in these cases,Constantwill not know whichQuantityto try converting to!
Otherwise — when you have a literal, and you’re passing it to a concrete Quantity API — either
one will work, and the choice is a matter of taste and local context. Consider the feature
comparison matrix above, and experiment and find out what works best for your codebase!
As always, if you have any feedback for us or encounter any problems, feel free to file an issue.
-
We needed this comment in our Aurora-internal implementations, for instance. ↩
-
Note that this is not true for Au. Our embedded teams had a seat at the table from the very beginning, and robust handling for integer-backed quantities is a core strength of Au. ↩
-
The “rep” is a common shorthand term in units libraries for “representation type”. It’s the underlying numeric type that stores the value of a particular quantity. ↩
-
This would be a line like
constexpr auto km = kilo(m);, in the same section at the top of the file where we import the symbols we use. Expand the “Includes and usings” section in the Au tab of our Eigen example for an example of this. ↩ -
Technically, we have not been able to strictly verify this claim, because Au does not include a definition for the “country mile” unit. 😁 ↩