Skip to content
Cole Page
ProjectHigh Yield

As Filed, Part Three

A progress update: moving to bond indentures, reading 808 of them with a model on my laptop, a mistake I made about the ones it got wrong, a small classifier experiment that didn't go the way I expected, and the first scores against labels I've checked myself.

Moog's 2014 indenture from start to finish, 2 minutes 13 seconds, no sound.

Part two ended with me saying the next post would be a hosted version of the app, and this isn't that, because sometime in the week after I published it I talked myself into the idea that I'd been doing things in the wrong order. Everything I'd built for covenants was built around bank credit agreements, and most of the companies I'd actually want to look at have bonds out as well, and bond covenants work differently enough that I don't think you can say you understand what a company is allowed to do until you've read both. So most of the last three weeks went into indentures, into getting the reading to run on my own laptop instead of paying for it, and into a couple of detours, one of which (a small classifier model) did something a bit different from what I thought it would. This is more of a where-things-stand update than a finished piece, everything in it is as of September 28, and a fair amount of it is still moving.

Why indentures are their own problem

The covenant I cared about in part two is a maintenance test, the kind where a borrower has to keep leverage under something like 4.50 to 1 and gets tested on it every quarter whether or not anything happens. A high-yield indenture mostly doesn't work like that. Its covenants only matter when the company tries to do something (borrow more, pay a dividend, put a lien on an asset, sell a business, get bought), and whether it's allowed to usually comes down to a ratio test, something like "only if fixed charge coverage would still be above 2.0 times", plus a stack of baskets that let it do specific amounts regardless of the test. So there isn't really a list of covenants to pull out the way there is in a loan agreement. There's one package per indenture, and the row I ended up with has the debt test and its ratio, the basket for bank debt, the restricted payments builder and general baskets, the liens test, the asset sale terms, the change of control put, the call schedule and the equity clawback, and every value has a quote and a character position behind it, same as the loan side.

There are also two kinds of document that don't fit that shape at all. Some indentures are written in the investment-grade style, where there's a negative pledge and not a lot else, and a lot of what gets filed is a supplemental indenture that just says the base indenture's covenants apply. The model has to be able to say either of those things plainly, and a surprising amount of the early work was getting it to do that instead of filling the fields with something plausible.

Counting the corpus again

Indentures show up on EDGAR as EX-4 exhibits, and the index I built has 125,014 of them. 5,627 carry what I've been calling the high-yield signature, which is roughly the vocabulary you'd expect (a fixed charge coverage test, restricted payments, the usual defined terms). I froze the text of 3,000 of those, from 941 companies filing between 2012 and 2026, and the set I actually ran is the latest indenture for each company, which came to 808 documents. The thinking there was that if you were looking at a company today, the newest indenture is the one you'd want first.

Reading them on a laptop

In part two everything went through a hosted model, and at about 53 cents an indenture (more than a loan agreement, because the answer has so many fields that each document gets split into three calls) the whole corpus was going to cost somewhere in the low thousands of dollars, which is more than I wanted to spend finding out whether this works. So I set up an open model on my laptop, a MacBook with 128 GB of memory (which turns out to matter a lot for this), tuned it on the same small set of documents I'd used to score the hosted model, and then let it run. The 808 took 77.5 hours of machine time over five days and cost nothing beyond electricity, and 793 of them came back with a package the validator accepted. It rejected the other 15.

To give you a feel for what one read is like, Moog's 2014 indenture (the one in the video further down) took five calls and six minutes and 21 seconds. Three of those calls were the model working through the sections it was handed, and two were second tries after something in its first answer didn't check out against the text. Each call was carrying about 27,500 tokens of indenture.

For a while the number I kept quoting was that the laptop model agrees with the hosted one about 86% of the time, on the eleven documents I have a reference label for. I was fairly happy with that until I ran the exact same laptop setup twice on the same twelve documents and found it only agrees with itself 84% of the time. Once you put intervals on those numbers, they overlap almost completely:

comparisonagreement95% intervaldocuments
laptop model against the hosted model86.0%75.9% to 92.9%11
the laptop model against itself, two identical runs84.1%74.5% to 91.8%12

So every comparison I'd been making on that small set was inside the noise, which was a slightly deflating afternoon, but I'd much rather know it. The fix is more documents rather than more runs, and that's what the 25-document holdout set is for, which I finally got labelled this week (more on that further down).

What went wrong, and a thing I had wrong

The failure that mattered most was the model saying the covenants weren't there at all. At the end of the first pass, 60 of the 793 packages were in a state I'd started calling unresolved, meaning the model said it hadn't been shown the incurrence covenants. When I dug in, most of those turned out to be routing problems rather than reading problems. The router is the piece that decides which sections of a 300,000-character document each call gets to see, and it was trusting the document outline's label for which section held the debt covenant. The outline was sometimes wrong (it would hand that label to the section about reports to holders, or the one about guarantees), so the real covenant never made it into the context. I fixed that, re-read the 60, and it came down to 30. Then a count I'd set up for a different reason turned up two plain bugs in the outline itself: a cross-reference to a section's own number would steal that section's place, and a heading that wraps onto a second line would lose everything after the first. Fixing those and re-reading the 23 documents it changed got it down to 20.

At that point I wrote in my notes that the last 20 were the documents whose covenant the outline still couldn't find, mostly older indentures that number their sections like 1008 instead of 10.08, and I believed that for about five days. It fell apart this week, when the classifier experiment below happened to show the router handing over the debt covenant on 9 of the 12 remaining documents that actually state one. I didn't trust that, so I went back and rebuilt the routing exactly as it was on the day those documents were read, and it was the same nine. The model had the operative sentence right there in its context, somewhere inside 55,000 to 214,000 characters of debt-related text, and it still said it hadn't been given one. What I'd actually done was notice that 20 unresolved documents was very close to the 19 that the outline couldn't reach, and treat two similar counts as if they were the same set of documents. So what's left is mostly the model not finding something it was given, which is a different problem, and the numbering fix I'd planned would help with three of them at most.

The plan all along has been that anything the laptop model gets wrong, or isn't sure about, goes to the hosted model for a second read. Right now that's 69 documents and roughly $37, and I haven't spent it yet, partly because the list keeps getting shorter every time I fix something upstream.

A detour into classifiers

There's been a lot of talk over the last couple of weeks about small classifier models, the BERT kind that were state of the art before large language models took over, coming back for the jobs they're better at, which are fast, cheap decisions about a piece of text where you don't need anything written back. Deciding which passages of an indenture describe which covenant is exactly that sort of decision, and it's the one my pipeline had been getting wrong, so I trained one. The labels were more or less free, since the 793 documents the laptop model had already read tell you which passages it quoted for each covenant. I split everything by company so it was never tested on an issuer it had seen, and I wrote down what I expected before looking at any results, which was that it would find the covenant on at least ten of the unresolved documents, where I thought the router had failed.

That turned out to be wrong in two ways. The router hadn't failed on most of them (that's how I found the mistake above), and the classifier found the covenant on 7 of the 12, against 9 for the router. They found different ones, though, and between them they cover 11 of the 12. And on the companies it never saw in training, it's very good at the narrower thing it was trained to do. Its top five passages for a covenant contain the text the reading model ended up quoting 99.6% of the time, against 96.5% for the router, while pointing at about 6,000 characters of text instead of the router's 54,000. It scores a whole indenture in about 12 seconds.

I'm not calling that a win yet, for two reasons. The labels are the laptop model's own quotes, which it made from text the router had shown it, so there's something a bit circular about asking whether the classifier can find them again. And the sentence a model ends up quoting isn't everything it needs to read, since so much of an indenture hangs off the definitions. The only real test is to read with it, so that's what the laptop is doing while I write this: sixteen documents read twice, once with the router's context and once with the classifier's much smaller one, and then the twelve unresolved documents the same way, to see whether a tighter context helps a model that was probably drowning in 200,000 characters.

The test finished while I was writing this, and it came out somewhere in the middle. Reading with the classifier's passages instead of the router's cut the text going into each call by 17 to 23% and the time a document by 10 to 15%, which is real but a long way from the ninefold saving I'd half expected, because most of what goes into a call (the instructions, the worked example, the definitions) stays the same no matter which passages you swap in. Quality is harder to call. Its agreement with the reference labels came out lower, 77% against 85% for the router, but when I looked at where the two differ, most of the gap is the classifier's version filling in fields the reference left empty (make-whole spreads, lien baskets, asset sale terms), almost always with a quote behind them, so I suspect it's partly a more complete read being marked against a less complete answer key, and I want to check those by hand before I believe either number. The part I found most interesting was the twelve unresolved documents. On a fresh read the router's version resolved 4 of them and the classifier's resolved 7, including 2 of the 3 the router had never shown the covenant at all. Twelve documents is too few to lean on, but if it holds up, the natural job for the classifier is a second try when the first read comes back empty.

The holdout, finally

The holdout is 25 indentures from companies that aren't in anything I tuned on, 5 of which I've set aside blind and won't score until I've settled on a setup. Labelling 25 indentures at the level of detail the package has is a lot of hours, so I did it the way I'd want to check anyone else's work. A large hosted model drafted every document twice, in two separate passes that worked only from the frozen text and never saw each other's answers, and wherever the two drafts agreed on a field I took it as settled. Where they disagreed (14 fields out of about 900) I read the filing and decided, and I also checked a random sample of the fields they'd agreed on, where they turned out to be wrong on 1 of the 18 I could check. I did the same for a new set of 20 credit agreements on the loan side.

The first time I scored the laptop model against those labels, on the 20 documents that aren't blind, it agreed on 70.6% of fields with the router's context and 71.8% with the classifier's, each give or take about six points. That's a lot lower than the 84 to 86% I'd been quoting, although those were two models agreeing with each other rather than a model agreeing with a person, and the labels turned out to be more complete than either model's read (most of the misses were fields the label filled in and the model left empty). The parts I care about most held up, with the type of debt test right on 19 of 20 documents and its ratio right on all 16 that have one, and change of control and the notes' own terms both around 90%. Two parts were badly broken, though. Liens were at 10 to 15% and redemption (the call schedule, the equity clawback and the make-whole spread) at 37 to 62%, and they were broken in the same way whichever routing I used, which usually means the problem is further upstream.

It turned out to be two things. The general lien basket in a high-yield indenture is usually one clause sitting deep inside the definition of Permitted Liens, and the router only passes along the first 3,000 characters of a definition, so on most documents the model never saw the number, and the make-whole spread had a similar problem, since it lives inside a definition the router wasn't prioritising. The classifier did show the model the basket (on 9 of the 10 documents that have one), and the model still answered that there wasn't one, because my own instructions described a plain negative pledge as what most high-yield indentures carry, and the one worked example I give it happens to be an indenture with no lien basket and no make-whole spread. So I'd written a rule that said one thing and an example that taught the other, and the model followed the example.

I fixed both in general terms (longer definitions for the two that matter, and instructions that agree with the rule) and read the holdout again. The classifier version went from 71.8% to 76.6%, with liens going from 15% to 40% and redemption from 62% to 69%, and when I ran it a second time on identical code to see how much of that was noise, the repeat landed within a point of it, at 77.3%. The router version went from 70.6% to 73.0%. This second read isn't a clean holdout number anymore, since the holdout is what told me what to fix, so I'm keeping the first result on the record beside it, and the blind five stay unscored until I've picked a setup. Liens still aren't good, though, and the model hasn't once said a lien covenant is keyed to a ratio, even on the seven documents where the label says it is (most of which also have a dollar basket, and my rule says the ratio wins), and I think the fix there is a field that can hold both, tried out on a small labelled set that isn't the holdout. On the classifier question, with the fixed code it's ahead of the router by 2.8 points, with an interval running from 3.3 behind to 9.3 ahead, so I still can't say it's better, only that it isn't worse and isn't slower.

I ran the loan side against its new holdout as well, 16 credit agreements the laptop model hadn't seen (four more are blind). The first read came in at 0.73 F1, well under the 0.87 from part two, although that was the hosted model on different documents, so it isn't a clean comparison. Part of the gap was one misreading. Some agreements list their covenants under "The Borrower will not" with each one starting "Permit the Total Leverage Ratio to exceed", and the model was reading that bare "Permit" as permission to do something rather than as the covenant itself. With that fixed, two identical reads came in at 0.87 and 0.83, so some of the improvement is the fix and some of it is how much a single run moves around on 16 documents, which is about four points between two runs that should in principle be the same.

The video

The video at the top of this post is one indenture, Moog's for its 5.250% notes due 2022, going through the whole pipeline. I made it mostly because it's a lot easier to show someone what this does than to explain it, and everything in it comes from the data. The 324,861 characters of text are the filing's, the sections are the outline's, the scores in the routing step are the classifier's real scores for that document (Moog was held out, so it never saw it in training), and every value lands on the exact characters it was quoted from. The clock during the reading step is real as well, the six minutes and 21 seconds that read took, running at thirty times speed. The one thing in it that isn't how the pipeline runs today is the routing step, since the classifier is still an experiment and the reads you see were made with the router's context. I'd rather show where this is heading and say so than show last month's version, and since it's generated from the data rather than animated by hand, I can regenerate it when the numbers move.

Where this is heading

The table behind all this has 806 indentures in it and 8,610 located citations, and none of it has been reviewed by a person yet, which is the part that's on me. The workbench from part two works for bonds now too, so it's really just a matter of sitting down with it. The routing test didn't turn up the big speedup I'd been hoping for, so getting through the other 2,200 or so indentures I've already frozen, and eventually the rest of the filings behind them, is going to keep the laptop busy for a while either way. The next things I'll try are the classifier as that second pass on the reads that come back empty and the lien field that can hold both a ratio and a basket, and after that it's loans again, the same way, and then the hosted app, which I'm now calling part four. The project is still private for the moment.