<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/">
    <channel>
        <title>Matt Goodman</title>
        <link>https://goodmattg.xyz</link>
        <description>Your blog description</description>
        <lastBuildDate>Sun, 16 Aug 2026 20:20:33 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>https://github.com/jpmonette/feed</generator>
        <image>
            <title>Matt Goodman</title>
            <url>https://goodmattg.xyz/favicon.ico</url>
            <link>https://goodmattg.xyz</link>
        </image>
        <copyright>All rights reserved 2026</copyright>
        <item>
            <title><![CDATA[Ashenbrenner Can Go F**** Himself]]></title>
            <link>https://goodmattg.xyz/articles/Ashenbrenner-Review</link>
            <guid>https://goodmattg.xyz/articles/Ashenbrenner-Review</guid>
            <pubDate>Fri, 15 Nov 2024 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<blockquote>
<p>Algorithmic progress: In the coming decade, AI labs will invest tens of billions in algorithmic R&amp;D, and all the smartest people in the world will be working on this; from tiny efficiencies to new paradigms, we'll be picking lots of the low-hanging fruit. We probably won't reach any sort of hard limit (though “unhobblings” are likely finite), but at the very least the pace of improvements should slow down, as the rapid growth (in $ and human capital investments) necessarily slows down (e.g., most of the smart STEM talent will already be working on AI). (That said, this is the most uncertain to predict, and the source of most of the uncertainty on the OOMs in the 2030s on the plot above.)</p>
</blockquote>
<p>I recently read Ashenbrenner's several hundred page missive about the coming inevitability of AGI filled with charts that go up and to the right, and my only thought was: “go f**** yourself”.  But then I realized that if he spent his time carefully interviewing people at OpenAI and writing hundreds of pages to lay out his argument, I should at least throw together one poorly researched page to rebut it.</p>
<p>My biggest issue is with his statement “most of the smart STEM talent will already be working on AI”. Wrong. Most people just do what the market has support for, and the market has limited support for people getting paid &gt;300k a year to hunt for algorithmic inefficiencies on which they WILL NOT MAKE MONEY. These are not high frequency algorithmic trading systems. The major returns from foundation model companies to this point have been from sale to large FAANG companies looking to bolster their own internal capabilities, NONE OF WHICH MAKE MEANINGFUL MONEY ON GENAI. “GENAI” does not currently make any money for anyone other than the Cloud Providers (OpenAI, etc.), the top 2-3 chat providers (ChatGPT, Claude, Characters, etc.), and the top 2-3 image generators (ChatGPT / Imagen, Adobe, StableDiffusion, etc.) - “AI” as a business enabler is far more than GENAI methods including large scale recommendation algorithms (Amazon, TikTok, Netflix, Youtube), RL systems (market makers, ad pricing systems ), and traditional “ML” / robotics systems (warehouses, drone systems, construction, logistics) to improve core business processes. All of Meta and Google's revenue comes from Ads. The core use of “AI” in those businesses is to improve their profit centers. Google's / Microsoft / Amazon's revenue from GENAI is a drop in the bucket on their P&amp;L, and should really be viewed as another line in the cloud services business.</p>
<p>I smell a grift. Let's look at base incentives. Forget the rhetoric for a second. Pretend you're an LLM / “AI” researcher - one of the real ones from the top CS programs (Stanford, Berkeley, MIT, CMU, Princeton, Oxford, etc). You likely started your academic career after AlexNet in 2012, and finished your doctorate sometime 2014-2024. If you finished your Ph.D. before that, you could easily be an esteemed professor who did the hard work of defining the field, and can still probably make money lending your name as an advisor to an AI company or  a dual appointment at a tech company, Let's say this pool of extremely elite people is ~5000 people (10 schools, 50 Phds awarded in relevant research areas, 10 years), but we all know that's generous. There's actually like ~1000 people who <em>really</em> define the field. We don't have to pretend someone with a Ph.D. in language theory, or geometric computer vision, distributed system, or hell, even cryptography, has made any real contribution to the LLM specific arms race we're in now.  No. It's guys like Sutskever, Bengio, Shamir, Hassabis and a few other people who either wrote the core research or joined OpenAI / Google Research / DeepMind early enough to catch the wave. Cool, great. These people, by the way are world-class academics, elite organization builders, and highly proficient engineers. Now hand them billions of dollars of investor capital and an army of research engineers that are the best at setting up large compute clusters, and getting models to concurrently run on thousands to tens of thousands (and more) GPU's concurrently using really really good distributed code.</p>
<blockquote>
<p>AI progress won't stop at human-level. Hundreds of millions of AGIs could automate AI research, compressing a decade of algorithmic progress (5+ OOMs) into 1 year. We would rapidly go from human-level to vastly superhuman AI systems. The power—and the peril—of superintelligence would be dramatic.</p>
</blockquote>
<p>Continuing the incentives talk, you're one of these people with all the knowledge, the keys to the kingdom in the form of equity in a highly valued LLM venture, and we're at the industry wide benchmarks of GPT-4o and Llama V2. These systems are amazing, but if Ashenbrenner is to believed, we're still on the exponentially rising curve that will end with AGI in 2030, so your equity in OpenAI / Alphabet / Nvidia is still exponentially undervalued with the AGI to come. Why would you sell any shares, or for that reason leave at all? You've got an equity stake in literally bringing our AGI overlords to fruition - surely that stake will be honored with piles of money by the 2030's “at the latest”.</p>
<p>My view is much more sober and, frankly, far less fun.</p>
<p><strong>The brilliant people who built these systems have seen the writing on the wall - that we can continue to improve these systems via exponential investment in compute, but there are a limited number of OOMs up for grabs and funding is finite because VC funds have ~5-7 years to deliver returns. Furthermore, the differentiation between the best models for the mean of knowledge work is declining - so pricing power and margins will decline as well. There will be a handful of winners that will win because of network effects and enterprise adoption, squeezing out upstarts that have to burn all their funding on compute, expensive researchers, and outmatched sales teams. Now is the time cash out equity and raise money to form new ventures while the AI investment thesis is still de jour and there is a still a premium placed on AI expertise.</strong></p>
<p>It's a perfectly hedged decision given the downside is rejoining a FAANG research group in 2-3 years.</p>
<p>Here's what I believe is happening at a blocking and tackling level. Collect a few brand name Ph.Ds who worked at any "world-class" LLM research organization that have some risk tolerance, split off to form a new company. Use the academic star power and “ex” mafia to raise &gt;100mm dollars, build a chat or GenAI offering aimed at enterprise in some knowledge work centered vertical (law, health, finance, advertising, marketing, etc), go to war with the other &gt;100mm funded enterprise AI companies for a finite pool of companies with the scale and means to use AI to improve their knowledge work processes but CRITICALLY not technologically sophisticated enough to either (1) deploy free open source models into their organization (e.g. Llama) or (2) be one of the ultra-profitable FAANG enterprises that won't use your AI tools anyways. Now play startup for awhile. If your company finds PMF, great! If it doesn't or you can't fundraise more because the economics don't support it (read - you aren't profitable), then get acquired / merge back into FAANG at a reasonable multiple.</p>
<p>We've already seen this quite a few times. Shamir left, went to Captions, now back at Google. Ilya left OpenAI, raised something like &gt;100mm to start SSI…. but we'll see. I'd be willing to bet he'll be back at one of Alphabet, Microsoft, Apple, or Nvidia within 5 years. A bunch of the best robotics and RL people left to go do robotics foundation models at Physical Intelligence, but unless they actually solve that and build a model that completely revolutionizes all future robotics problems (which is unlikely), they'll be rehired as the DeepMind robotics research group within 5 years. Eric Jang is doing this at X1, and Fei-Fei Li (et. Al) are doing this at [I FORGET]. At some point Mistral will fold into a European conglomerate. Anthropic is essentially funded by Amazon and will eventually fold back in.</p>
<p><strong>This time is not different, we're just seeing the circle of life.</strong></p>
<p>Venture capital is de-facto funding the best academics to keep doing their core research, but at wildly inflated multiples. And for the the thousands of lesser known companies building “AI for X”, the question is whether their tools will so drastically improve the productivity / revenue of their target market that they can get pricing power. If you're selling me AI for CRM management, and I only had 1 person beforehand doing that job, I can now fire that person and improve my profitability via the productivity improvement, but will it completely revolutionize my business to where I'll pay your SaaS fees? Probably not. I don't buy that this productivity improvement will lead to a new economy, or even wholesale change many industries, it will just shuffle the productivity formula of knowledge work, causing some low-skill knowledge jobs to be replaced by software, and new knowledge jobs emerge. Meanwhile industries that actually employ large numbers of people (healthcare, manufacturing, trade labor) will be unaffected in the near term.</p>
<p>All of this to get back to why Ashenbreener can go f**** himself; he can show me a hundred curves that are up and to the right, but those curves are descriptive - none of them tell me how an AI system, which is just a piece of software running on machines, will start to run itself. Because at the end of the day, even the most sophisticated LLM “AI” models we have now are just wildly effective seq2seq models. We've managed to compress <em>all</em> previously written  (and soon visual) knowledge implicitly into a function parameterized by trillions of numbers and begun injecting that system into our business operations, but it's still just a system without a self. Anyone who speaks this confidently on something deserves to be checked - because  think about it, if you were actually this CERTAIN that AGI is coming by 2030, shouldn't you quit your job and move to  to await the end times with the other doomsday cults? What is the point of life and building these systems if you truly believe the end is INEVITABLE. I personally don't think Ashenbrenner believes it either, because he has the gall to lay out all the heavy statements about inevitability but the cowardice to use phrasing like “probably”, "would", and "could" before every statement. Well sir, I won't be bullied by “charts” and “math” into accepting our AGI overlords, and neither should you. Show me your trades Ashenbrenner - you coward, and go f**** yourself.</p>
<p>-- Goodman</p>
]]></content:encoded>
            <author>gmgprivacy@proton.me (Matt Goodman)</author>
        </item>
        <item>
            <title><![CDATA[Thoughts on Consuming and Making]]></title>
            <link>https://goodmattg.xyz/articles/Consuming</link>
            <guid>https://goodmattg.xyz/articles/Consuming</guid>
            <pubDate>Fri, 21 Jun 2024 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Consumers are N=1 Makers</p>
<p>Don’t be a consumer, it’s doesn’t lead to a life worth living.</p>
<p>We hear all the time about the strength of the American Consumer, or the consumer economy, but what actually is a consumer? Not in the obvious “a consumer consumes” sense - a consumer is someone who buys things of course, but how we label our economy as driven by consumers. When you start looking, you can find lots of wannabe-VC-god-king dichotomies of people in the world, but the Consumer vs. Creator dichotomy is real - there are people who make things: products, companies, content, and then there are people who buy them. The buyers vastly outnumber the sellers, which is how you get an economic pyramid that concentrates wealth in the hands of a few individuals. Take a second to think of all the things you buy and consume: services, products, food, media. Do you make any of these?</p>
<p>From what I’ve seen, most people “make” only one thing - they do their job for their employer, “making” economic value. They then consume a salary for that work which they funnel into more consumption of all the stuff I described above. For every influencer that has some unbelievably rich 5-9 after they’re 9-5, most people just go to work, maybe exercise, then watch TikTok / Netflix / media, then go to bed. The whole day they consume food and media and services, and they only put 1 thing into the world.</p>
<p>The Consumer is what capitalism pushes us to be, consumers of everything around us up to and beyond how much money we have, consuming experiences and new restaurants and taking the same trips that we’ve seen online to the same places with the same photos. I speak from personal experience, there is a hollowness to being a Consumer - it’s akin to gluttony or hedonism, you just have to fill your role, consume, and then die.</p>
<p>So my advice is to go N &gt; 1. What does that mean? It means rewiring your brain and daily routines. Our brains make it incredibly easy to fall into a pattern of joy through consumption - pulling up TikTok for 30 minutes leads to a nice dopamine wave, but you get nothing from it in the end. Writing an essay (like this one) that will probably be bad and no one will ever read is hard and doesn’t make you happy most of the time, but you made a thing, and you can point at it and say “I made a thing”. Making a thing whether it’s art or music or software or a meal worth sharing is the whole point. That probably means you’ll be a less traditionally "fun" person because you’ll start slipping on contacts and social media references, but you have to make choices, and after a median lifetime of consuming, you want to drastically over-correct.</p>]]></content:encoded>
            <author>gmgprivacy@proton.me (Matt Goodman)</author>
        </item>
        <item>
            <title><![CDATA[The Demand for Crypto: Smart People Want Lucrative Jobs]]></title>
            <link>https://goodmattg.xyz/articles/Demand-for-Crypto</link>
            <guid>https://goodmattg.xyz/articles/Demand-for-Crypto</guid>
            <pubDate>Sun, 09 May 2021 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<h2>The Demand: smart people who want lucrative jobs correlated with "smartness"</h2>
<p>I <em>still</em> don't get crypto, but I kind of get one of the needs for crypto. If you run in my "<em>elite</em>" <sup><a href="#user-content-fn-1" id="user-content-fnref-1" data-footnote-ref="true" aria-describedby="footnote-label">1</a></sup> post-grad circle you'll notice that a lot of people talk  about finance. Every quantitatively oriented software engineer I meet will introduce themselves and be like "I'm John. In my free time I like hiking, hanging out with friends, and I do some quantitative trading <strong>just for fun</strong>." Just for fun?! Is quantitative trading fun? I mean, it's lucrative if you have good strategies and the right infrastructure, but is it actually so "fun" that we would do it in our unpaid free time? What I'm getting at is lots of people in the academic <em>elite</em> that know quant trading is a way to (1) be smart (2) make gobs of money where (3) the amount of money you make is correlated with how smart you are. And if you consider the kinds of people who make it to <em>elite</em> institutions, this is the dream job. Most of us learn young that how much you earn isn't tied to how smart you are, but this is tough pill for people who <em>won</em> at "being smart". You spend a lifetime jumping through academic hoops, winning the race, and you get to the end where statistically, it hardly matters <sup><a href="#user-content-fn-2" id="user-content-fnref-2" data-footnote-ref="true" aria-describedby="footnote-label">2</a></sup>. But then, there's this one magical niche field, quantitative trading, where the two strongly correlate, so naturally the demand to enter the field is high. I studied electrical engineering, but after looking at the EE job market as a rational economic actor, I realized joining the field would be silly. If you love EE and would die doing it in your free time, sure go work in EE. But if you're me and only really liked EE, you do what it takes to switch towards a field with higher compensation. Why the low EE salaries? I'm not an economist, but rising income inequality, international labor markets, and skilled worker visas (H-1B) successfully commoditized the EE profession. So in the end, EE and other hard science / engineering careers pay well, but not nearly the money software engineering and quant finance pay.</p>
<h1>The Supply: exceptional organizations that don't need many people</h1>
<p>But from the quantitative finance side <sup><a href="#user-content-fn-3" id="user-content-fnref-3" data-footnote-ref="true" aria-describedby="footnote-label">3</a></sup>, the demand for new talent is far lower than supply. What is this statement based on? Anecdotal evidence, so call foul if I'm off. But just eyeballing my peer group, it looks like Jane Street / etc. take a handful of new graduates a year maximum. Assume a small multiple above new grads for experienced hires. Any of these places employ at maximum ~2000 people, and it's pretty clear these small exceptional organizations minimize their headcount so everyone gets paid more. Nothing surprising here. And so, we kind of see that quant finance shops always need new talent, because people flame out or retire, but not nearly as much talent as <em>elite</em> institutions produce. From personal conversations, it feels like &gt;50% of computer science, statistics, math, and physics Ph.Ds. have considered quant finance.</p>
<h1>The Mismatch</h1>
<p>There's a mismatch here. Lot's of highly intelligent, credentialed people that could probably do these jobs, not many jobs because the firms are massively scalable. Citadel Securities reports they account for 26% of all US equities volume with ~1400 employees <sup><a href="#user-content-fn-4" id="user-content-fnref-4" data-footnote-ref="true" aria-describedby="footnote-label">4</a></sup>. The larger quant hedge funds have less than 1k employees, and Renaissance famously employs like what, 300 people? Quant finance is just software and finance IP - these firms don't produce tangible goods and stay lean to maximize profitability. And worse, more market participants means more efficient markets - i.e. less alpha. Again, un-cited, but this why IP is so closely guarded at these firms. Strategies disappear as soon as other market participants can execute the same strategy. It's reductive, but I just think of all these guys as alpha miners. Finite amount of gold on Earth, finite amount of alpha in the markets. Unless we could just create more markets....</p>
<h1>Cryptocurrencies = Infinite "Amateur Golden Windows"</h1>
<p>So here you go: I think below the surface, some of the demand for crypto is just masked demand for new markets. We need crypto because we need more markets. We need more markets because there are a lot of smart people who want to correlate their earnings with their intelligence. Alpha decreases monotonically in existing markets, so the solution is new markets. Cryptocurrencies present an <em>elegant</em> way to create new FX instruments.</p>
<p><strong>Q</strong>: What's a challenge with FX trading from the growth perspective?<br>
<strong>A</strong>: Currencies are 1:1 with nation-states, and the world isn't getting any new nation-states.</p>
<br>
<br>
<p><strong>Q</strong>: But why can't we make money trading traditional FX?<br>
<strong>A</strong>: You <em>can</em> make money, but there's a difference between betting on directions/movements (gambling) and systematic trading. FX is the largest and most liquid market in the world. With so many market participants and massive liquidity, you can bet it's pretty goddamn difficult to find alpha, let alone arbitrages.</p>
<br>
<br>
<p><strong>Q</strong>: So what makes crypto different?<br>
<strong>A</strong>: From one perspective, nothing. Every new cryptocurrency just creates <em>N</em> new currency pairs with <em>N</em> existing currencies. But because the cryptos have different flavors / features / auto-generated copy, you get weird cross currency dynamics. <strong>Figuring out weird market dynamics with algorithms is mining alpha</strong>.</p>
<br>
<br>
<p><strong>Q</strong>: Couldn't we create markets with other asset classes?<br>
<strong>A</strong>: Of course, but isn't it easier to go with something where you just click a button and a new instrument springs into existence globally?</p>
<br>
<br>
<p><strong>Q</strong>: So how is this a solution for people who don't already work professionally in <em>traditional</em> quantitative finance?<br>
<strong>A</strong>: There's a <strong>golden window of time</strong> before trading volume draws in professionals where amateurs can rediscover existing strategies and apply them new instruments. You can bet that as <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>t</mi><mo>→</mo><mi mathvariant="normal">∞</mi></mrow><annotation encoding="application/x-tex">t \rightarrow \infty</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6151em"></span><span class="mord mathnormal">t</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">→</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.4306em"></span><span class="mord">∞</span></span></span></span>, the big HFT shops will deploy the same cross-exchange latency arbitrage strategies they use now, but those strategies either can't be guaranteed to work or aren't worth the cost to deploy until:</p>
<ol>
<li>Instrument trading volumes are large</li>
<li>Market data infrastructure is stable</li>
<li>There is enough market data to backtest new strategies</li>
</ol>
<p>To hammer it home, cryptocurrencies are in one way valuable because they provide an <strong>infinite number of golden windows</strong> for amateurs to exploit before volumes draw in bigger players. By the time bigger players enter the market, some kid can hit a button and create a new market on an existing blockchain, or a group of engineers can write a new blockchain. Rinse and repeat. Your smart friends can trade crypto and conceivably make money that isn't just random gambling. Larger quant finance institutions support this because they get to deploy large-scale / high-volume strategies and take the alpha if any cryptocurrencies become widely successful (e.g. Bitcoin, Ether). Are cryptocurrencies "good"? Unlikely. Are cryptocurrencies valuable? They are if people value them.</p>
<section data-footnotes="true" class="footnotes"><h2 class="sr-only" id="footnote-label">Footnotes</h2>
<ol>
<li id="user-content-fn-1">
<p>I mean this in the most self-deprecating way possible, it's just faster than writing "Ivy League + all the elite technical schools, elite state schools, honors programs at state schools, anyone who worked hard at any educational institution, and brilliant people who took alternative paths". <a href="#user-content-fnref-1" data-footnote-backref="" aria-label="Back to reference 1" class="data-footnote-backref">↩</a></p>
</li>
<li id="user-content-fn-2">
<p>...relative to other college educated persons. The earnings gap by education level is well demonstrated in economics literature. <a href="#user-content-fnref-2" data-footnote-backref="" aria-label="Back to reference 2" class="data-footnote-backref">↩</a></p>
</li>
<li id="user-content-fn-3">
<p>To save some time I'm going to lump all quantitative finance into a bucket, but yes, the field has many sub-disciplines: HFT, prop trading shops, quant hedge funds, quant groups at bulge brackets, etc. <a href="#user-content-fnref-3" data-footnote-backref="" aria-label="Back to reference 3" class="data-footnote-backref">↩</a></p>
</li>
<li id="user-content-fn-4">
<p><a href="https://www.citadelsecurities.com/products/equities-and-options/">https://www.citadelsecurities.com/products/equities-and-options/</a> <a href="#user-content-fnref-4" data-footnote-backref="" aria-label="Back to reference 4" class="data-footnote-backref">↩</a></p>
</li>
</ol>
</section>]]></content:encoded>
            <author>gmgprivacy@proton.me (Matt Goodman)</author>
        </item>
        <item>
            <title><![CDATA[Dynamic Time Warping for Clustering Time Series Data]]></title>
            <link>https://goodmattg.xyz/articles/Dynamic-Time-Warping</link>
            <guid>https://goodmattg.xyz/articles/Dynamic-Time-Warping</guid>
            <pubDate>Sun, 10 Dec 2017 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<h2>Motivation</h2>
<p>This was for the Wikipedia competition hosted by Kaggle competition to predict future page view counts on Wikipedia. Like so many of the solutions now the winning approach used RNNs to learn a set of hyperparameters on the entire dataset that are then fed in to a predict locally on each time series. This begs comparison to multivel hierarchical models from a Bayesian perspective, but I'll save that for a later post. I treat these competitions as a way to explore new classes of techniques more than as competition to win. And it's from that perspective that I settled on Dynamic Time Warping as the technique I wanted to explore. The big question I had was:</p>
<blockquote>
<p>Instead of learning hyperparameters on our entire corpus of time series, wouldn't it be great if we could find a way to sub divide our corpus into highly <strong>similar</strong> groups, and then learn hyperparameters on those groups?</p>
</blockquote>
<p>In the context of the [winning solution]<sup><a href="#user-content-fn-1" id="user-content-fnref-1" data-footnote-ref="true" aria-describedby="footnote-label">1</a></sup> this clustering layer would sit a layer above the RNN. In the language of hierarchical multilevel models, this would adding another level to the model that where we then compute posterior distributions on the hyperparameters for each group, with additional hyperparameters for the entire population. This yielded the question of how to cluster time series data.
I looked to the research and settled on Dynamic Time Warping (DTW) from Keogh et. al from UC Riverside. For any pair of time series signals, we can find the energy of the signal that minimizes the distance between two input signals. DTW can easily show that time warped signals are the same. This technique has been highly effective in classifying EEG and other time-varying signals. As</p>
<h2>Literature Review: DTW</h2>
<p>This review is informal but should provide sufficient grounding. Most of the important work has been done by Keogh</p>
<ol>
<li>
<p>Rakthanmanon, T., Campana, B., Mueen, A., Batista, G., Westover, B., Zhu, Q., … Keogh, E. (n.d.). Searching and Mining Trillions of Time Series Subsequences under Dynamic Time Warping.</p>
</li>
<li>
<p>Mueen, A., &amp; Keogh, E. (2016). Extracting Optimal Performance from Dynamic Time Warping, 2129–2130.</p>
</li>
<li>
<p>Lemire, D. (2009). Faster retrieval with a two-pass dynamic-time-warping lower bound, 42, 2169–2180. <a href="http://doi.org/10.1016/j.patcog.2008.11.030">http://doi.org/10.1016/j.patcog.2008.11.030</a></p>
</li>
<li>
<p>Keogh, E. (n.d.). Clustering of Time Series Subsequences is Meaningless : Implications for Previous and Future Research.</p>
</li>
<li>
<p>H. Sakoe and S. Chiba, “Dynamic programming algorithm optimization for spoken word recognition,” IEEE Trans. Acoust. Speech, Lang. Process., vol. 26, no. 1, pp. 43–50, 1978.</p>
</li>
<li>
<p>MuÌller, Meinard. Information Retrieval for Music and Motion. Springer, 2010.</p>
</li>
</ol>
<h2>Theory: DTW</h2>
<p>For an excellent review of the computational aspects of DTW I recommend "Abdullah Mueen, Eamonn J. Keogh: Extracting Optimal Performance from Dynamic Time Warping. KDD 2016: 2129-2130". Refer to Meinard for a solid review of DTW theory.</p>
<p>Dynamic Time Warping is a path-searching algorithm. DTW finds the minimum cost path between the complete matrix of pairwise distances between two time-series we will label <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>X</mi></mrow><annotation encoding="application/x-tex">X</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07847em">X</span></span></span></span> and <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>Y</mi></mrow><annotation encoding="application/x-tex">Y</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.22222em">Y</span></span></span></span>. Define <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>X</mi><mo>:</mo><mo>=</mo><mo stretchy="false">(</mo><msub><mi>x</mi><mn>1</mn></msub><mo separator="true">,</mo><msub><mi>x</mi><mn>2</mn></msub><mo separator="true">,</mo><mi mathvariant="normal">.</mi><mi mathvariant="normal">.</mi><mi mathvariant="normal">.</mi><mo separator="true">,</mo><msub><mi>x</mi><mi>n</mi></msub><mo stretchy="false">)</mo></mrow><annotation encoding="application/x-tex">X := (x_1, x_2, ..., x_n)</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07847em">X</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">:=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mopen">(</span><span class="mord"><span class="mord mathnormal">x</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3011em"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight">1</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em"><span></span></span></span></span></span></span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord"><span class="mord mathnormal">x</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3011em"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight">2</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em"><span></span></span></span></span></span></span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord">...</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord"><span class="mord mathnormal">x</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.1514em"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">n</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em"><span></span></span></span></span></span></span><span class="mclose">)</span></span></span></span> and <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>Y</mi><mo>:</mo><mo>=</mo><mo stretchy="false">(</mo><msub><mi>y</mi><mn>1</mn></msub><mo separator="true">,</mo><msub><mi>y</mi><mn>2</mn></msub><mo separator="true">,</mo><mi mathvariant="normal">.</mi><mi mathvariant="normal">.</mi><mi mathvariant="normal">.</mi><mo separator="true">,</mo><msub><mi>y</mi><mi>n</mi></msub><mo stretchy="false">)</mo></mrow><annotation encoding="application/x-tex">Y:=(y_1, y_2,...,y_n)</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.22222em">Y</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">:=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mopen">(</span><span class="mord"><span class="mord mathnormal" style="margin-right:0.03588em">y</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3011em"><span style="top:-2.55em;margin-left:-0.0359em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight">1</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em"><span></span></span></span></span></span></span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.03588em">y</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3011em"><span style="top:-2.55em;margin-left:-0.0359em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight">2</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em"><span></span></span></span></span></span></span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord">...</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.03588em">y</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.1514em"><span style="top:-2.55em;margin-left:-0.0359em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">n</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em"><span></span></span></span></span></span></span><span class="mclose">)</span></span></span></span>. This matrix of pairwise distances is referred to as the <em>cost</em> matrix <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>C</mi></mrow><annotation encoding="application/x-tex">C</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07153em">C</span></span></span></span>. Define the function <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>c</mi><mo stretchy="false">(</mo><mi>x</mi><mo separator="true">,</mo><mi>y</mi><mo stretchy="false">)</mo></mrow><annotation encoding="application/x-tex">c(x,y)</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mord mathnormal">c</span><span class="mopen">(</span><span class="mord mathnormal">x</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord mathnormal" style="margin-right:0.03588em">y</span><span class="mclose">)</span></span></span></span> as the cost or local distance function between two points <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>x</mi></mrow><annotation encoding="application/x-tex">x</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em"></span><span class="mord mathnormal">x</span></span></span></span> and <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>y</mi></mrow><annotation encoding="application/x-tex">y</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.625em;vertical-align:-0.1944em"></span><span class="mord mathnormal" style="margin-right:0.03588em">y</span></span></span></span>. If <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>c</mi><mo stretchy="false">(</mo><mi>x</mi><mo separator="true">,</mo><mi>y</mi><mo stretchy="false">)</mo><mo>=</mo><mn>0</mn></mrow><annotation encoding="application/x-tex">c(x,y)=0</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mord mathnormal">c</span><span class="mopen">(</span><span class="mord mathnormal">x</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord mathnormal" style="margin-right:0.03588em">y</span><span class="mclose">)</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.6444em"></span><span class="mord">0</span></span></span></span>, the two points are identical. Therefore, low cost implies similarity, high cost implies dissimilarity. DTW finds a path through the cost matrix of minimum total cost. Each valid path through the cost matrix is called a "warping" path. The set of all warping paths is labeled <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>W</mi></mrow><annotation encoding="application/x-tex">W</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.13889em">W</span></span></span></span>.</p>
<p>This recursive function gives the minimum cost path:</p>
<p><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>γ</mi><mo stretchy="false">(</mo><mi>i</mi><mo separator="true">,</mo><mi>j</mi><mo stretchy="false">)</mo><mtext>&nbsp;</mtext><mo>=</mo><mtext>&nbsp;</mtext><mi>d</mi><mo stretchy="false">(</mo><msub><mi>q</mi><mi>i</mi></msub><mo separator="true">,</mo><msub><mi>c</mi><mi>j</mi></msub><mo stretchy="false">)</mo><mo>+</mo><mi>m</mi><mi>i</mi><mi>n</mi><mo stretchy="false">{</mo><mi>γ</mi><mo stretchy="false">(</mo><mi>i</mi><mo>−</mo><mn>1</mn><mo separator="true">,</mo><mi>j</mi><mo>−</mo><mn>1</mn><mo stretchy="false">)</mo><mo separator="true">,</mo><mi>γ</mi><mo stretchy="false">(</mo><mi>i</mi><mo>−</mo><mn>1</mn><mo separator="true">,</mo><mi>j</mi><mo stretchy="false">)</mo><mo separator="true">,</mo><mi>γ</mi><mo stretchy="false">(</mo><mi>i</mi><mo separator="true">,</mo><mi>j</mi><mo>−</mo><mn>1</mn><mo stretchy="false">)</mo><mo stretchy="false">}</mo></mrow><annotation encoding="application/x-tex">\gamma(i,j)~=~d(q_i, c_j)+min\{\gamma(i-1,j-1), \gamma(i-1, j), \gamma(i, j-1)\}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1.0361em;vertical-align:-0.2861em"></span><span class="mord mathnormal" style="margin-right:0.05556em">γ</span><span class="mopen">(</span><span class="mord mathnormal">i</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord mathnormal" style="margin-right:0.05724em">j</span><span class="mclose">)</span><span class="mspace nobreak">&nbsp;</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace nobreak">&nbsp;</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mord mathnormal">d</span><span class="mopen">(</span><span class="mord"><span class="mord mathnormal" style="margin-right:0.03588em">q</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em"><span style="top:-2.55em;margin-left:-0.0359em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">i</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em"><span></span></span></span></span></span></span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord"><span class="mord mathnormal">c</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight" style="margin-right:0.05724em">j</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2861em"><span></span></span></span></span></span></span><span class="mclose">)</span><span class="mspace" style="margin-right:0.2222em"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mord mathnormal">min</span><span class="mopen">{</span><span class="mord mathnormal" style="margin-right:0.05556em">γ</span><span class="mopen">(</span><span class="mord mathnormal">i</span><span class="mspace" style="margin-right:0.2222em"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em"></span></span><span class="base"><span class="strut" style="height:0.854em;vertical-align:-0.1944em"></span><span class="mord">1</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord mathnormal" style="margin-right:0.05724em">j</span><span class="mspace" style="margin-right:0.2222em"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mord">1</span><span class="mclose">)</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord mathnormal" style="margin-right:0.05556em">γ</span><span class="mopen">(</span><span class="mord mathnormal">i</span><span class="mspace" style="margin-right:0.2222em"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mord">1</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord mathnormal" style="margin-right:0.05724em">j</span><span class="mclose">)</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord mathnormal" style="margin-right:0.05556em">γ</span><span class="mopen">(</span><span class="mord mathnormal">i</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord mathnormal" style="margin-right:0.05724em">j</span><span class="mspace" style="margin-right:0.2222em"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mord">1</span><span class="mclose">)}</span></span></span></span></p>
<p><img src="/assets/posts/DTW/DTWTrace.jpg" alt="DTW Trace">
<em>Figure 1. DTW computed optimal path between sinusoid &amp; sinusoid + sinc over <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo stretchy="false">[</mo><mn>0</mn><mo separator="true">,</mo><mn>6</mn><mi>π</mi><mo stretchy="false">]</mo></mrow><annotation encoding="application/x-tex">[0,6\pi]</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mopen">[</span><span class="mord">0</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord">6</span><span class="mord mathnormal" style="margin-right:0.03588em">π</span><span class="mclose">]</span></span></span></span></em></p>
<h2>DTW Nearest-Neigbhor Clustering</h2>
<p>The Wikipedia challenge provide 145,000 time-series to predict. Ideally, I would have computed pairwise DTW distances for all 145,000 pairs, and then run agglomerative clustering with Ward linkage to generate clusters. The computation wasn't feasible on my desktop setup, and I wasn't willing to go the AWS route for this exercise. Instead, I decided to randomly sample 10,000 rows, compute pairwise DTWs for the row sample, run hierarchical clustering on the pairwise distance matrix, and finally assign all non-sampled rows to the cluster matching of their nearest neighbor. Admittedly this was not the ideal solution. By assigning all non-sampled rows to the cluster containing their nearest neighbor (i.e. minimizing DTW distance), I likely unbalanced the clusters. Furthermore, nearest-neighbor clustering is meaningless with high-dimensionality inputs. I considered nearest-neighbor defensible because my dimensionality was unitary. That is, I only clustered on the DTW distance and so avoided have high-dimensionality inputs. I didn't compute any rough heuristics to quantify the distribution of non-sampled rows to clusters due to time constraints.</p>
<blockquote>
<p>"As noted by (Agrawal et al., 1993, Bradley and Fayyad, 1998), in high dimensions the very concept of nearest neighbor has little meaning, because the ratio of the distance to the nearest neighbor over the distance to the average neighbor rapidly approaches one as the dimensionality increases." [2]</p>
</blockquote>
<p>I ending up doing a minor rewrite of Prof. Eamonn Keogh's UCR-DTW C++ tool to do nearest-neighbor search instead of sub-sequence search. The tool now searches a list of time-series to find the one that minimizes the DTW distance instead of sub-sequence search to find the sub-sequence that minimizes distance. Because of optimizations in the UCR-DTW tool (via research by Lemire, Kim, et. al) the tool only computes the full DTW in 10% of cases.</p>
<p>After clustering, about 80,000 rows belong to one cluster - so we could conceptualize that cluster as the base data set and run our standard model. I would be curious to train the winning solution on each of my clusters to see if that improved its performance. The smaller cluster had predictable features like single large spikes, uniformity, two spikes, clear seasonality, etc. Using different linkage whole and partial linkage patterns instead of ward only increased the size of the largest cluster. This indicates that there is no signicant distinction for the majority of the time series signals in the corpus.</p>
<h2>Code Appendix</h2>
<pre class="language-matlab"><code class="language-matlab">
    <span class="token comment">%% Define original signal</span>
    x <span class="token operator">=</span> <span class="token number">0</span><span class="token operator">:</span><span class="token keyword">pi</span><span class="token operator">/</span><span class="token number">64</span><span class="token operator">:</span><span class="token number">6</span><span class="token operator">*</span><span class="token keyword">pi</span><span class="token punctuation">;</span>
    y <span class="token operator">=</span> <span class="token number">0.25</span><span class="token operator">*</span><span class="token function">sin</span><span class="token punctuation">(</span>x<span class="token punctuation">)</span><span class="token punctuation">;</span>

    <span class="token comment">%% Modify a section of the signal</span>
    N <span class="token operator">=</span> <span class="token function">length</span><span class="token punctuation">(</span>y<span class="token punctuation">)</span><span class="token punctuation">;</span>
    xind <span class="token operator">=</span> <span class="token punctuation">(</span><span class="token function">floor</span><span class="token punctuation">(</span><span class="token number">0.40</span><span class="token operator">*</span>N<span class="token punctuation">)</span><span class="token operator">:</span><span class="token function">floor</span><span class="token punctuation">(</span><span class="token number">0.60</span><span class="token operator">*</span>N<span class="token punctuation">)</span><span class="token punctuation">)</span><span class="token punctuation">;</span>
    ymod <span class="token operator">=</span> y<span class="token punctuation">;</span>
    <span class="token function">ymod</span><span class="token punctuation">(</span>xind<span class="token punctuation">)</span> <span class="token operator">=</span> <span class="token function">sinc</span><span class="token punctuation">(</span><span class="token function">x</span><span class="token punctuation">(</span>xind<span class="token punctuation">)</span><span class="token punctuation">)</span><span class="token punctuation">;</span>

    <span class="token comment">%% Compute DTW</span>
    zg <span class="token operator">=</span> <span class="token function">repmat</span><span class="token punctuation">(</span>y<span class="token punctuation">,</span> N<span class="token punctuation">,</span> <span class="token number">1</span><span class="token punctuation">)</span><span class="token punctuation">;</span>
    zgmod <span class="token operator">=</span> <span class="token function">repmat</span><span class="token punctuation">(</span>ymod<span class="token operator">'</span><span class="token punctuation">,</span> <span class="token number">1</span><span class="token punctuation">,</span> N<span class="token punctuation">)</span><span class="token punctuation">;</span>
    d2 <span class="token operator">=</span> <span class="token function">sqrt</span><span class="token punctuation">(</span><span class="token punctuation">(</span>zg <span class="token operator">-</span> zgmod<span class="token punctuation">)</span> <span class="token operator">.^</span> <span class="token number">2</span><span class="token punctuation">)</span><span class="token punctuation">;</span>
    <span class="token punctuation">[</span><span class="token operator">~</span><span class="token punctuation">,</span>ix<span class="token punctuation">,</span>iy<span class="token punctuation">]</span> <span class="token operator">=</span> <span class="token function">dtw</span><span class="token punctuation">(</span>y<span class="token punctuation">,</span> ymod<span class="token punctuation">)</span><span class="token punctuation">;</span>

    <span class="token comment">%% Plottingm</span>
    <span class="token punctuation">[</span>xg<span class="token punctuation">,</span> yg<span class="token punctuation">]</span> <span class="token operator">=</span> <span class="token function">meshgrid</span><span class="token punctuation">(</span><span class="token number">1</span><span class="token operator">:</span>N<span class="token punctuation">)</span><span class="token punctuation">;</span>
    <span class="token function">d2</span><span class="token punctuation">(</span><span class="token function">sub2ind</span><span class="token punctuation">(</span><span class="token function">size</span><span class="token punctuation">(</span>d2<span class="token punctuation">)</span><span class="token punctuation">,</span> iy<span class="token punctuation">,</span> ix<span class="token punctuation">)</span><span class="token punctuation">)</span> <span class="token operator">=</span> <span class="token number">1.5</span><span class="token operator">*</span><span class="token function">max</span><span class="token punctuation">(</span><span class="token function">d2</span><span class="token punctuation">(</span><span class="token operator">:</span><span class="token punctuation">)</span><span class="token punctuation">)</span><span class="token punctuation">;</span>
    <span class="token function">mesh</span><span class="token punctuation">(</span>xg<span class="token punctuation">,</span> yg<span class="token punctuation">,</span> d2<span class="token punctuation">)</span>
</code></pre>
<section data-footnotes="true" class="footnotes"><h2 class="sr-only" id="footnote-label">Footnotes</h2>
<ol>
<li id="user-content-fn-1">
<p><a href="https://www.kaggle.com/c/web-traffic-time-series-forecasting/discussion/43795">https://www.kaggle.com/c/web-traffic-time-series-forecasting/discussion/43795</a> <a href="#user-content-fnref-1" data-footnote-backref="" aria-label="Back to reference 1" class="data-footnote-backref">↩</a></p>
</li>
</ol>
</section>]]></content:encoded>
            <author>gmgprivacy@proton.me (Matt Goodman)</author>
        </item>
        <item>
            <title><![CDATA[Fable 5: The First Hit is Free]]></title>
            <link>https://goodmattg.xyz/articles/Fable-5-The-First-Hit-Is-Free</link>
            <guid>https://goodmattg.xyz/articles/Fable-5-The-First-Hit-Is-Free</guid>
            <pubDate>Fri, 12 Jun 2026 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<figure class="not-prose my-8"><img alt="Artemis II lifts off atop the SLS rocket, climbing skyward on a column of flame." loading="lazy" width="2400" height="959" decoding="async" data-nimg="1" style="color:transparent;background-size:cover;background-position:50% 50%;background-repeat:no-repeat;background-image:url(&quot;data:image/svg+xml;charset=utf-8,%3Csvg xmlns='http://www.w3.org/2000/svg' viewBox='0 0 320 120'%3E%3Cfilter id='b' color-interpolation-filters='sRGB'%3E%3CfeGaussianBlur stdDeviation='20'/%3E%3CfeColorMatrix values='1 0 0 0 0 0 1 0 0 0 0 0 1 0 0 0 0 0 100 -1' result='s'/%3E%3CfeFlood x='0' y='0' width='100%25' height='100%25'/%3E%3CfeComposite operator='out' in='s'/%3E%3CfeComposite in2='SourceGraphic'/%3E%3CfeGaussianBlur stdDeviation='20'/%3E%3C/filter%3E%3Cimage width='100%25' height='100%25' x='0' y='0' preserveAspectRatio='none' style='filter: url(%23b);' href='data:image/jpeg;base64,/9j/4AAQSkZJRgABAgAAAQABAAD/wAARCAADAAgDAREAAhEBAxEB/9sAQwAKBwcIBwYKCAgICwoKCw4YEA4NDQ4dFRYRGCMfJSQiHyIhJis3LyYpNCkhIjBBMTQ5Oz4+PiUuRElDPEg3PT47/9sAQwEKCwsODQ4cEBAcOygiKDs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7/8QAHwAAAQUBAQEBAQEAAAAAAAAAAAECAwQFBgcICQoL/8QAtRAAAgEDAwIEAwUFBAQAAAF9AQIDAAQRBRIhMUEGE1FhByJxFDKBkaEII0KxwRVS0fAkM2JyggkKFhcYGRolJicoKSo0NTY3ODk6Q0RFRkdISUpTVFVWV1hZWmNkZWZnaGlqc3R1dnd4eXqDhIWGh4iJipKTlJWWl5iZmqKjpKWmp6ipqrKztLW2t7i5usLDxMXGx8jJytLT1NXW19jZ2uHi4+Tl5ufo6erx8vP09fb3+Pn6/8QAHwEAAwEBAQEBAQEBAQAAAAAAAAECAwQFBgcICQoL/8QAtREAAgECBAQDBAcFBAQAAQJ3AAECAxEEBSExBhJBUQdhcRMiMoEIFEKRobHBCSMzUvAVYnLRChYkNOEl8RcYGRomJygpKjU2Nzg5OkNERUZHSElKU1RVVldYWVpjZGVmZ2hpanN0dXZ3eHl6goOEhYaHiImKkpOUlZaXmJmaoqOkpaanqKmqsrO0tba3uLm6wsPExcbHyMnK0tPU1dbX2Nna4uPk5ebn6Onq8vP09fb3+Pn6/9oADAMBAAIRAxEAPwDNljRtMeRlyyuAD6DBpxb5hSS5D//Z'/%3E%3C/svg%3E&quot;)" srcset="/_next/image?url=%2F_next%2Fstatic%2Fmedia%2Fartemis-ii-launch.8fbc3bf9.jpg&amp;w=3840&amp;q=75&amp;dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW 1x" src="/_next/image?url=%2F_next%2Fstatic%2Fmedia%2Fartemis-ii-launch.8fbc3bf9.jpg&amp;w=3840&amp;q=75&amp;dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW"><figcaption class="mt-3 text-sm leading-6 text-zinc-900 dark:text-zinc-100"><p>Artemis II lifts off from Launch Complex 39B, April 1, 2026. Image credit: NASA/Michael DeMocker.</p></figcaption></figure>
<p>I’ve been extensively testing <a href="https://www.anthropic.com/news/claude-fable-5-mythos-5">Fable 5 (Mythos class)</a> after the recent launch. My tests pushed its capabilities from spatial reasoning, critical reasoning, long running research and feedback, multi-step tasks, cinematography / visual design, and other more minor tasks.</p>
<p>Examples:</p>
<ul>
<li>help research, design, and procure CNC milled parts</li>
<li>assist with optics procurement</li>
<li>do cinematography for my company landing page hero (<a href="https://vigilautonomy.com">vigilautonomy.com</a>)</li>
</ul>
<p>My TLDR feedback is that without a doubt Anthropic cooked with this one. Using <a href="https://www.anthropic.com/news/claude-opus-4-8">Opus 4.8</a> or its competitor <a href="https://openai.com/index/introducing-gpt-5-5/">GPT 5.5 xhigh</a> always felt like working with a brilliant intern in their first month - it would do EXACTLY what you said because it was afraid of screwing up and wanted to please you. So if you put a typo in your prompt by gosh it would do its damndest to execute that typo.</p>
<p>On the one hand that’s great? Like I want this AI to be aligned with my intended use, so if I tell it something crazy and I meant it, then I don’t want pushback; just do it. But like most people I make typos all the time, and when I try to use voice to text transcription for technical problems those typos pop up even more often. The ideal alignment here is judgement of my prompt with reasoning to align with the intentions of my words, not the exact words themselves.</p>
<h2>Usage Patterns</h2>
<p>I do all of my tasks with permissions completely removed - mainly because I think that pattern is the future and the rewards absolutely outweigh the risks. Consider the idea of placing a human (me) in the middle of a loop of operations that may seem inscrutable, so when Fable is asking if I’m okay with an operation like ripgrep or curl or some chained together shell operation my only real reason to interject is not because I disagree with the METHOD, because I do not even attempt to understand the method, it’s because I’m scared of the blast radius and privilege escalation.</p>
<p>I fully believe that as we get more computer use data from Fable 5 users and the flywheel for the frontier labs spins up that post-training will be able to build in better guardrails on the system that make it harder to blow things up. But honestly, the consequences fall on me and I find them worth it to get the full agentic “go off and do this” experience so I can focus on other work.</p>
<p>The only best practice I always engage in is specifying allowed tools in my CLAUDE.md / AGENTS.md that I consider “safe” (chrome browser, ffmpeg, FreeCAD, python virtualenv, isolated worktrees), with explicit instructions to halt immediately if any execution ever requires changing GPU drivers or establishing persistent long running background tasks that will remain after prompt execution. I know the “safety” this brings me is untrue in so many ways, but in my scoped usage it has been effective. I think as long as you don’t say unbounded tasks like “go build me a 1 billion dollar business”, you’re going to be fine. <a href="https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/">METR</a> is a decent mental heuristic here - as long as the task I’m asking it to do is scopable within 1 day of human work, it’s likely the tool usage is contained enough.<sup><a href="#fn1">1</a></sup></p>
<p>For example I’ve already seen Fable 5 taking pains to limit its own blast radius even when I give it my personal thinkmaxxing cheat code of “exhaustively”:</p>
<ul>
<li>clone my code into an isolated worktree</li>
<li>spin up an ephemeral and completely separate python virtual environment for the debugging</li>
<li>use playwright to control an isolated chrome browser</li>
<li>all of this gets wiped as soon as the task is done</li>
</ul>
<p>Within the standard bounds of “use my computer to do a thing that is not crazy in a moderately well bounded filesystem with good enough available tools” Fable nails it every time.</p>
<h2>Case Study 1: Research, CAD, Procure</h2>
<p>I gave Fable the task of producing CAD artifacts of M6 adapter plates for my work at <a href="https://vigilautonomy.com">Vigil Autonomy</a> I could send to a machine on demand vendor that would be accepted without any issues, based on hardware casing specifications buried in PDFs and across websites. Consider a few of the steps involved if you were doing this manually:</p>
<ul>
<li>for each part go find the relevant spec document. If you’re lucky this is a STEP file, if unlucky it’s a technical drawing buried in a PDF</li>
<li>Translate that spec to a CAD file of an adapter plate that fits to an M6 optical breadboard</li>
<li>The CAD artifacts must be acceptance ready for a print on demand vendor (<a href="https://www.xometry.com">Xometry</a>, <a href="https://sendcutsend.com">SendCutSend</a>) and must use language a human reviewer would understand including acceptable technical drawing PDFs</li>
<li>Actual uploading of the CAD files to those vendor sites using menu usage to specify taps at each hole for the CNC mill</li>
<li>Procurement of any mounting equipment from a separate vendor like <a href="https://www.mcmaster.com">McMaster-Carr</a></li>
</ul>
<p>Put another way this is a complex chain of:</p>
<ul>
<li>OCR / Diagram Extraction</li>
<li>Tool use (CAD)</li>
<li>Judgement / Reasoning: Do the CAD files we designed make sense? Will they mount to the optical breadboard? Not only does the math check out, but do we have underflush? When we’re operating in field in Kuwait will the heat cause warping?</li>
<li>Search (for best practices working with vendors)</li>
<li>Tool Use + Judgement (Browser): now that we have a browser and a vendor site, how do we upload our “correct” CAD files in a way the vendor accepts? How do we examine other vendor sites to identify SKUs that will allow one-shot mounting?</li>
</ul>
<figure class="not-prose my-8"><img alt="Isometric CAD render of the Ellipse-D breadboard adapter plate showing corner bolt holes, M3 taps, and counterbored dowel bores." loading="lazy" width="1800" height="1800" decoding="async" data-nimg="1" class="border border-zinc-200 dark:border-zinc-800" style="color:transparent;background-size:cover;background-position:50% 50%;background-repeat:no-repeat;background-image:url(&quot;data:image/svg+xml;charset=utf-8,%3Csvg xmlns='http://www.w3.org/2000/svg' viewBox='0 0 320 320'%3E%3Cfilter id='b' color-interpolation-filters='sRGB'%3E%3CfeGaussianBlur stdDeviation='20'/%3E%3CfeColorMatrix values='1 0 0 0 0 0 1 0 0 0 0 0 1 0 0 0 0 0 100 -1' result='s'/%3E%3CfeFlood x='0' y='0' width='100%25' height='100%25'/%3E%3CfeComposite operator='out' in='s'/%3E%3CfeComposite in2='SourceGraphic'/%3E%3CfeGaussianBlur stdDeviation='20'/%3E%3C/filter%3E%3Cimage width='100%25' height='100%25' x='0' y='0' preserveAspectRatio='none' style='filter: url(%23b);' href='data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAAgAAAAICAYAAADED76LAAAAmUlEQVR42oVP6QqCQBjc93+SIDq8M1syMrCwogfo2OigYsUj/bHGTiUkSEQDw3x8zMAMwR+QzxEnKdYbBh5GKIpH3RDFCUbeFKpJYdhDTPw5tuyA9J6BvBPu2EdH65fs6k6l1PVAZsESbdVCS7FealfUrQGOpzOIEAKX6w3BYgWjR9FoKtBMp/zVSkopkeU5dmwPzsPvFb/wBK2y6fqN97RqAAAAAElFTkSuQmCC'/%3E%3C/svg%3E&quot;)" srcset="/_next/image?url=%2F_next%2Fstatic%2Fmedia%2Fellipse-d-adapter-render.104bbd6e.png&amp;w=1920&amp;q=75&amp;dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW 1x, /_next/image?url=%2F_next%2Fstatic%2Fmedia%2Fellipse-d-adapter-render.104bbd6e.png&amp;w=3840&amp;q=75&amp;dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW 2x" src="/_next/image?url=%2F_next%2Fstatic%2Fmedia%2Fellipse-d-adapter-render.104bbd6e.png&amp;w=3840&amp;q=75&amp;dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW"><img alt="Vendor-ready technical drawing of the Ellipse-D adapter plate with hole table, tolerances, and machining notes." loading="lazy" width="2420" height="1870" decoding="async" data-nimg="1" class="mt-4 border border-zinc-200 dark:border-zinc-800" style="color:transparent;background-size:cover;background-position:50% 50%;background-repeat:no-repeat;background-image:url(&quot;data:image/svg+xml;charset=utf-8,%3Csvg xmlns='http://www.w3.org/2000/svg' viewBox='0 0 320 240'%3E%3Cfilter id='b' color-interpolation-filters='sRGB'%3E%3CfeGaussianBlur stdDeviation='20'/%3E%3CfeColorMatrix values='1 0 0 0 0 0 1 0 0 0 0 0 1 0 0 0 0 0 100 -1' result='s'/%3E%3CfeFlood x='0' y='0' width='100%25' height='100%25'/%3E%3CfeComposite operator='out' in='s'/%3E%3CfeComposite in2='SourceGraphic'/%3E%3CfeGaussianBlur stdDeviation='20'/%3E%3C/filter%3E%3Cimage width='100%25' height='100%25' x='0' y='0' preserveAspectRatio='none' style='filter: url(%23b);' href='data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAAgAAAAGCAIAAABxZ0isAAAAWUlEQVR42lWMSwrAIAxEvf8tXYgbEX9oPi10bLDQt5rkJeNU9T4QUYwxpdR7dyJiDmGt5b0PIdRat7gO9mT5J/Cacy6lbIFBXuaczIw2BCzddwsgWmtjDBQ+wsaLWqiq6dAAAAAASUVORK5CYII='/%3E%3C/svg%3E&quot;)" srcset="/_next/image?url=%2F_next%2Fstatic%2Fmedia%2Fellipse-d-adapter-drawing.c24a3747.png&amp;w=3840&amp;q=75&amp;dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW 1x" src="/_next/image?url=%2F_next%2Fstatic%2Fmedia%2Fellipse-d-adapter-drawing.c24a3747.png&amp;w=3840&amp;q=75&amp;dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW"><figcaption class="mt-3 text-sm leading-6 text-zinc-900 dark:text-zinc-100"><p>One of the Fable 5 deliverables: the Ellipse-D breadboard adapter plate. CAD render (top) and the vendor-ready technical drawing (bottom) — 6061-T6, M3 taps, H7 dowel bores, callouts a machinist reviewer can act on. <a href="/files/ellipse-d-adapter-drawing.pdf" class="underline">Download the drawing PDF</a>.</p></figcaption></figure>
<h3>Analysis</h3>
<p>We’ll only know after mounting and field use, but Fable’s planning and execution matched the exact process I would have taken, accomplishing in 3 hours of execution with intermittent feedback to push for continued use in producing more CAD files, and the file synthesis into the browser what would have taken me 3-4 days of work?!</p>
<p>Obviously adapter plates are the ideal solid you could ask a model to generate - they are rigid, bounded, no deformations, single standardized material, and would require a human’s limited mental context to maintain (taps, flush, proud, underflush). But still…. holy shit.</p>
<p>When I asked it to produce validation diagrams showing its understanding of layout on the breadboard, how the screws would work on the board with threading distance, underflush available, it nailed it. When I asked it to confirm that our designs would succeed assuming the tolerances of a 3rd party machinist, Fable explained that obviously it had already considered this, surfaced its thinking tokens to prove it, and made it clear that any error here would be because the machinist did not match their quoted tolerance. Time will tell.</p>
<p>What I love the most is that Fable finally feels like that intern that’s been there for awhile - it’s still slavishly devoted to your goals, but when you’re wrong in your prompts it will reason about your actual goal, do that, and explain that it knew what you meant. I finally had zero times where it executed based on a typo (if I type “produce a STEO cad file” it knows I mean STEP).</p>
<p>I’m guessing as part of their post-training Anthropic is pushing more reasoning tokens during planning before any execution takes place. This makes a ton of sense - for safety aligned AI you need to know the INTENTION of someone’s prompt and filter red team prompts that may read as innocuous but are part of a malicious prompt chain. If someone asks for a list of chemical reagents or electronics related to explosives, reasoning would make a normal person ask “wait, why do you need this? I’m suspicious”, and as future prompts arrive I have that suspicion tied to reasoning, and can apply judgement to stay aligned with my safety policy.</p>
<p>By actively blocking usage for biotech / cybersecurity applications Anthropic is aggressively enforcing their safety mandate<sup><a href="#fn2">2</a></sup> - which I tend to agree with, and the downstream effect of all the upfront reasoning is a model that nails intention alignment.</p>
<p>The multi-step execution for long running flows is solved. The 1M token contexts are more than enough memory for my isolated workflows, Anthropic has clearly made improvements in context efficiency, and when in doubt memory files are the right solution for picking up from an earlier checkpoint. There was no task I gave it, including the cinematography tasks for <a href="https://vigilautonomy.com">vigilautonomy.com</a>, where it failed, we’re talking long running ffmpeg processes, wiring videos into the landing page, and visual analysis of other best practices landing sites.</p>
<h2>Case Study 2: Hero Video Production</h2>
<p>Until I hire professional designers I’ve been stress testing LLM visual reasoning by repeatedly redesigning the landing page for my company Vigil Autonomy (<a href="https://vigilautonomy.com">vigilautonomy.com</a> - hit me up if you like aerial autonomy). LLMs are crushing all of the technical benchmarks but we all know they still don’t nail aesthetic taste.</p>
<p>I think this will continue to be true for a few reasons: (1) design inspiration may be from the distribution (i.e. “inspiration” or “reference”), but new designs are inherently out of distribution (2) design language is ambiguous (what is “clean” design, describe “blue” to a blind person) and (3) LLMs are not inherently human - the designs they produce cannot link to a gut feeling (the Golden Arches suggest nothing in the transistors of an LLM).</p>
<p>I gave it the task of generating a split-screen piece of cinematography I could use in the hero. On the left I had some UAV footage taken during a test flight from above, on the right I wanted paired footage from a ground station detecting the UAV.</p>
<p>You get into interesting problems - on the left it’s more about visually inspiring and cinematic shots with “clean” clip transitions (see even I can’t help use ambiguous language) - on the right I wanted my ground station UAV capture footage, but I didn’t just want hover shots of the drone, I wanted to capture the actual importance of why this data matters - the drone flying from out of frame or a distance towards the camera. I had a loose idea of the clips that looked nice, but didn’t want to do any labeling of frames or help with trajectory segment selection.</p>
<p>Here’s what Fable 5 did:</p>
<ul>
<li>first planned how it would accomplish a split-screen design with ffmpeg</li>
<li>then did clip selection from my ground footage by collating the known UAV trajectories, where the UAV would land in frame in post projection, filtering segments down to make sure even after cropping we’d see the UAV fly into the split screen frame</li>
<li>executed long running FFmpeg background subprocesses</li>
<li>updated my landing page to use the video</li>
</ul>
<figure class="not-prose my-8"><img alt="Hand-drawn storyboard sketch of the split-screen hero: cinematic footage on the left, data capture with stitched clips and nanosecond timestamps on the right." loading="lazy" width="1600" height="1200" decoding="async" data-nimg="1" style="color:transparent;background-size:cover;background-position:50% 50%;background-repeat:no-repeat;background-image:url(&quot;data:image/svg+xml;charset=utf-8,%3Csvg xmlns='http://www.w3.org/2000/svg' viewBox='0 0 320 240'%3E%3Cfilter id='b' color-interpolation-filters='sRGB'%3E%3CfeGaussianBlur stdDeviation='20'/%3E%3CfeColorMatrix values='1 0 0 0 0 0 1 0 0 0 0 0 1 0 0 0 0 0 100 -1' result='s'/%3E%3CfeFlood x='0' y='0' width='100%25' height='100%25'/%3E%3CfeComposite operator='out' in='s'/%3E%3CfeComposite in2='SourceGraphic'/%3E%3CfeGaussianBlur stdDeviation='20'/%3E%3C/filter%3E%3Cimage width='100%25' height='100%25' x='0' y='0' preserveAspectRatio='none' style='filter: url(%23b);' href='data:image/jpeg;base64,/9j/4AAQSkZJRgABAgAAAQABAAD/wAARCAAGAAgDAREAAhEBAxEB/9sAQwAKBwcIBwYKCAgICwoKCw4YEA4NDQ4dFRYRGCMfJSQiHyIhJis3LyYpNCkhIjBBMTQ5Oz4+PiUuRElDPEg3PT47/9sAQwEKCwsODQ4cEBAcOygiKDs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7/8QAHwAAAQUBAQEBAQEAAAAAAAAAAAECAwQFBgcICQoL/8QAtRAAAgEDAwIEAwUFBAQAAAF9AQIDAAQRBRIhMUEGE1FhByJxFDKBkaEII0KxwRVS0fAkM2JyggkKFhcYGRolJicoKSo0NTY3ODk6Q0RFRkdISUpTVFVWV1hZWmNkZWZnaGlqc3R1dnd4eXqDhIWGh4iJipKTlJWWl5iZmqKjpKWmp6ipqrKztLW2t7i5usLDxMXGx8jJytLT1NXW19jZ2uHi4+Tl5ufo6erx8vP09fb3+Pn6/8QAHwEAAwEBAQEBAQEBAQAAAAAAAAECAwQFBgcICQoL/8QAtREAAgECBAQDBAcFBAQAAQJ3AAECAxEEBSExBhJBUQdhcRMiMoEIFEKRobHBCSMzUvAVYnLRChYkNOEl8RcYGRomJygpKjU2Nzg5OkNERUZHSElKU1RVVldYWVpjZGVmZ2hpanN0dXZ3eHl6goOEhYaHiImKkpOUlZaXmJmaoqOkpaanqKmqsrO0tba3uLm6wsPExcbHyMnK0tPU1dbX2Nna4uPk5ebn6Onq8vP09fb3+Pn6/9oADAMBAAIRAxEAPwCe1eWS9UF+N2etZMpH/9k='/%3E%3C/svg%3E&quot;)" srcset="/_next/image?url=%2F_next%2Fstatic%2Fmedia%2Fsplit-screen-storyboard.da1a5749.jpg&amp;w=1920&amp;q=75&amp;dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW 1x, /_next/image?url=%2F_next%2Fstatic%2Fmedia%2Fsplit-screen-storyboard.da1a5749.jpg&amp;w=3840&amp;q=75&amp;dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW 2x" src="/_next/image?url=%2F_next%2Fstatic%2Fmedia%2Fsplit-screen-storyboard.da1a5749.jpg&amp;w=3840&amp;q=75&amp;dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW"><figcaption class="mt-3 text-sm leading-6 text-zinc-900 dark:text-zinc-100"><p>The entire design spec I gave Fable 5: a notebook sketch of the split-screen concept — cinematic on the left, data on the right.</p></figcaption></figure>
<figure class="not-prose my-8"><img alt="Static frame of the finished split-screen hero billboard: cinematic aerial footage over Austin on the left, ground-station capture of the UAV flying in over a highway on the right." loading="lazy" width="2400" height="675" decoding="async" data-nimg="1" style="color:transparent;background-size:cover;background-position:50% 50%;background-repeat:no-repeat;background-image:url(&quot;data:image/svg+xml;charset=utf-8,%3Csvg xmlns='http://www.w3.org/2000/svg' viewBox='0 0 320 80'%3E%3Cfilter id='b' color-interpolation-filters='sRGB'%3E%3CfeGaussianBlur stdDeviation='20'/%3E%3CfeColorMatrix values='1 0 0 0 0 0 1 0 0 0 0 0 1 0 0 0 0 0 100 -1' result='s'/%3E%3CfeFlood x='0' y='0' width='100%25' height='100%25'/%3E%3CfeComposite operator='out' in='s'/%3E%3CfeComposite in2='SourceGraphic'/%3E%3CfeGaussianBlur stdDeviation='20'/%3E%3C/filter%3E%3Cimage width='100%25' height='100%25' x='0' y='0' preserveAspectRatio='none' style='filter: url(%23b);' href='data:image/jpeg;base64,/9j/4AAQSkZJRgABAgAAAQABAAD/wAARCAACAAgDAREAAhEBAxEB/9sAQwAKBwcIBwYKCAgICwoKCw4YEA4NDQ4dFRYRGCMfJSQiHyIhJis3LyYpNCkhIjBBMTQ5Oz4+PiUuRElDPEg3PT47/9sAQwEKCwsODQ4cEBAcOygiKDs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7Ozs7/8QAHwAAAQUBAQEBAQEAAAAAAAAAAAECAwQFBgcICQoL/8QAtRAAAgEDAwIEAwUFBAQAAAF9AQIDAAQRBRIhMUEGE1FhByJxFDKBkaEII0KxwRVS0fAkM2JyggkKFhcYGRolJicoKSo0NTY3ODk6Q0RFRkdISUpTVFVWV1hZWmNkZWZnaGlqc3R1dnd4eXqDhIWGh4iJipKTlJWWl5iZmqKjpKWmp6ipqrKztLW2t7i5usLDxMXGx8jJytLT1NXW19jZ2uHi4+Tl5ufo6erx8vP09fb3+Pn6/8QAHwEAAwEBAQEBAQEBAQAAAAAAAAECAwQFBgcICQoL/8QAtREAAgECBAQDBAcFBAQAAQJ3AAECAxEEBSExBhJBUQdhcRMiMoEIFEKRobHBCSMzUvAVYnLRChYkNOEl8RcYGRomJygpKjU2Nzg5OkNERUZHSElKU1RVVldYWVpjZGVmZ2hpanN0dXZ3eHl6goOEhYaHiImKkpOUlZaXmJmaoqOkpaanqKmqsrO0tba3uLm6wsPExcbHyMnK0tPU1dbX2Nna4uPk5ebn6Onq8vP09fb3+Pn6/9oADAMBAAIRAxEAPwDndQvbsXchF1NnJ/5aGs5NrqKCT3P/2Q=='/%3E%3C/svg%3E&quot;)" srcset="/_next/image?url=%2F_next%2Fstatic%2Fmedia%2Fhero-billboard.74dd1f98.jpg&amp;w=3840&amp;q=75&amp;dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW 1x" src="/_next/image?url=%2F_next%2Fstatic%2Fmedia%2Fhero-billboard.74dd1f98.jpg&amp;w=3840&amp;q=75&amp;dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW"><figcaption class="mt-3 text-sm leading-6 text-zinc-900 dark:text-zinc-100"><p>The finished hero billboard on vigilautonomy.com — cinematic UAV footage left, ground-station detection footage right, angled seam per the sketch.</p></figcaption></figure>
<figure class="not-prose my-8"><img alt="The Vigil UI rerun view for a trajectory: a 3D scene with the UAV trajectory path and rover position alongside synchronized ground-station camera streams." loading="lazy" width="3200" height="2000" decoding="async" data-nimg="1" class="border border-zinc-200 dark:border-zinc-800" style="color:transparent;background-size:cover;background-position:50% 50%;background-repeat:no-repeat;background-image:url(&quot;data:image/svg+xml;charset=utf-8,%3Csvg xmlns='http://www.w3.org/2000/svg' viewBox='0 0 320 200'%3E%3Cfilter id='b' color-interpolation-filters='sRGB'%3E%3CfeGaussianBlur stdDeviation='20'/%3E%3CfeColorMatrix values='1 0 0 0 0 0 1 0 0 0 0 0 1 0 0 0 0 0 100 -1' result='s'/%3E%3CfeFlood x='0' y='0' width='100%25' height='100%25'/%3E%3CfeComposite operator='out' in='s'/%3E%3CfeComposite in2='SourceGraphic'/%3E%3CfeGaussianBlur stdDeviation='20'/%3E%3C/filter%3E%3Cimage width='100%25' height='100%25' x='0' y='0' preserveAspectRatio='none' style='filter: url(%23b);' href='data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAAgAAAAFCAIAAAD38zoCAAAAiElEQVR42gF9AIL/ABccHiEmJBodGhkaFx0eHjM0NTc5OjM0NgARGRgjNyQrOyYxOSZWWU7X2dbh4eKvsbEADRUTJDohLkAjNj4kU1VJxMfFz9DQoqOjAA0VEyY6HjA/IDg/IllbTd3h3err67e4uQANFRYjMSEqMh0tMR5KTEOkqKesra2EhYWRbSl+9HuMcgAAAABJRU5ErkJggg=='/%3E%3C/svg%3E&quot;)" srcset="/_next/image?url=%2F_next%2Fstatic%2Fmedia%2Fvigil-rerun-trajectory.aa6800dc.png&amp;w=3840&amp;q=75&amp;dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW 1x" src="/_next/image?url=%2F_next%2Fstatic%2Fmedia%2Fvigil-rerun-trajectory.aa6800dc.png&amp;w=3840&amp;q=75&amp;dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW"><figcaption class="mt-3 text-sm leading-6 text-zinc-900 dark:text-zinc-100"><p>The Vigil UI rerun view for a captured trajectory — the 3D trajectory path with synchronized ground-station camera streams.</p></figcaption></figure>
<h3>Analysis</h3>
<p>I don’t know much about cinematography (or if that’s even the right word) but as a computer vision person I’ve spent enough time looking at video to know this looks better than average. Like is it cinema grade? No. Is it better than what I had before? Yes. I liked that Fable 5 was able to handle the FFmpeg usage to do fairly obvious things like clipping / splicing that have always been painful. This was a series of annoying manual steps that are finally in reach because of a tool like this.</p>
<h2>Final Thoughts</h2>
<p>I’m going to need a better way to work around or with the robots.txt and agents.txt policies more websites are using and that Anthropic respects.<sup><a href="#fn3">3</a></sup> My intended usage is not malicious - if I want to go buy a list of things from <a href="https://www.mcmaster.com">McMaster-Carr</a>, I need to make sure I have the full up-to-date list of SKUs and specs so Fable 5 can reason, but the policies block this (fairly, IMO).</p>
<p>The tradeoff here will be how much I personally care, the future plan is to just have a less scrupulous agent (local LLM) go execute the scrape, build the cache, RAG it, then feed that local file to Fable 5 as it executes. Less token usage and we have the data we need so I can speed run my purchases. I’m dead certain Anthropic / OpenAI / Google / Stripe are fighting to sell this integration for agentic commerce into vendors like this, but in the meantime I still need to accelerate.</p>
<p>Finally, using a model like Fable 5 with high, xhigh, or max effort makes me as the operator wonder if I’m being ambitious enough. It feels like if I can come up with a very long complicated task with a healthy amount of detail the LLM will nail it - so I’m now bounded by big problems I can come up with. This is a joy - drudgery is now outsourced and the hard physical problems that need solving can get my entire focus without sacrificing anything.</p>
<p>Per my <a href="/articles/Mythos-and-DeFi">earlier article</a> the conclusion is this feels like a step change even if that change is just because the model is 25% better - it has surpassed some threshold. I’m not excited for this to move behind API pricing as it’s currently worth it, but at usage based pricing I’ll have to re-evaluate. I may find it worth it to use it as a coordinator leaving execution to dumber models, but we’re talking about more scaffolding, more mental overhead, more complexity. As they say, the first hit is free.</p>
<hr>
<h3>Footnotes</h3>
<ol>
<li><span id="fn1"></span> METR, <a href="https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/">“Measuring AI Ability to Complete Long Tasks”</a> (<a href="https://arxiv.org/abs/2503.14499">arXiv:2503.14499</a>) — the 50%-task-completion time horizon metric. The heuristic: if a skilled human could finish the task in about a day, the agent’s tool use tends to stay well bounded.</li>
<li><span id="fn2"></span> Anthropic, <a href="https://www.anthropic.com/news/claude-fable-5-mythos-5">“Introducing Claude Fable 5 and Claude Mythos 5”</a> — Fable 5 and Mythos 5 share the same underlying model; Fable 5 ships with additional safety measures for dual-use capabilities, while Mythos 5 is available only to approved organizations.</li>
<li><span id="fn3"></span> Anthropic, <a href="https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler">“Does Anthropic crawl data from the web, and how can site owners block the crawler?”</a> — ClaudeBot, Claude-User, and Claude-SearchBot honor robots.txt directives and anti-circumvention technologies.</li>
</ol>
<div class="not-prose mt-16 font-mono text-xs uppercase tracking-widest text-zinc-300 dark:text-zinc-700"><p>Edited with Fable 5</p></div>]]></content:encoded>
            <author>gmgprivacy@proton.me (Matt Goodman)</author>
        </item>
        <item>
            <title><![CDATA[Deep Learning in 2021]]></title>
            <link>https://goodmattg.xyz/articles/How-To-Train-A-Neural-Network-In-2021</link>
            <guid>https://goodmattg.xyz/articles/How-To-Train-A-Neural-Network-In-2021</guid>
            <pubDate>Sun, 09 May 2021 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>This post is going to cover the state of deep learning in 2021 with particular emphasis on AI infrastructure and deep learning tooling. If you're coming from a theory-based university classroom, get ready for an exhausting amount of detail. The classic pattern of "just feed some data through the network and backpropogate" is still the truth, but it takes ~10x more effort beyond the network to get anything useful.</p>
<h2>Theory</h2>
<p>TODO</p>
<h2>Requirements</h2>
<h3>Language: Python</h3>
<p>The lingua franca of modern DL is Python. All the major research codebases are in Python, the major DL frameworks all have Python bindings, and the vast majority of tooling is written for Python users. I don't know of any big league research organizations that use anything other than Python.</p>
<h2>CPU Training on Machine</h2>
<p>This is the "Hello World" of DL, but we'll explore this scenario in some detail so we can reference it in later sections. Let's use the <code>torchvision</code> library from the PyTorch team to train a simple classifier on MNIST.</p>
<pre class="language-python"><code class="language-python">
<span class="token keyword">import</span> torch<span class="token punctuation">.</span>nn <span class="token keyword">as</span> nn
<span class="token keyword">import</span> torch<span class="token punctuation">.</span>nn<span class="token punctuation">.</span>functional <span class="token keyword">as</span> F
<span class="token keyword">import</span> torch<span class="token punctuation">.</span>optim <span class="token keyword">as</span> optim
<span class="token keyword">import</span> torchvision<span class="token punctuation">.</span>transforms <span class="token keyword">as</span> transforms

<span class="token keyword">from</span> torch<span class="token punctuation">.</span>utils<span class="token punctuation">.</span>data <span class="token keyword">import</span> DataLoader
<span class="token keyword">from</span> torchvision<span class="token punctuation">.</span>datasets <span class="token keyword">import</span> MNIST

transform <span class="token operator">=</span> transforms<span class="token punctuation">.</span>Compose<span class="token punctuation">(</span><span class="token punctuation">[</span>
                               transforms<span class="token punctuation">.</span>ToTensor<span class="token punctuation">(</span><span class="token punctuation">)</span><span class="token punctuation">,</span>
                               transforms<span class="token punctuation">.</span>Normalize<span class="token punctuation">(</span>
                                 <span class="token punctuation">(</span><span class="token number">0.1307</span><span class="token punctuation">,</span><span class="token punctuation">)</span><span class="token punctuation">,</span> <span class="token punctuation">(</span><span class="token number">0.3081</span><span class="token punctuation">,</span><span class="token punctuation">)</span><span class="token punctuation">)</span>
                             <span class="token punctuation">]</span><span class="token punctuation">)</span>

<span class="token comment"># NOTE: Dataset downloaded to local machine</span>
trainset <span class="token operator">=</span> MNIST<span class="token punctuation">(</span>root<span class="token operator">=</span><span class="token string">'./data'</span><span class="token punctuation">,</span> train<span class="token operator">=</span><span class="token boolean">True</span><span class="token punctuation">,</span> download<span class="token operator">=</span><span class="token boolean">True</span><span class="token punctuation">,</span> transform<span class="token operator">=</span>transform<span class="token punctuation">)</span>
trainloader <span class="token operator">=</span> DataLoader<span class="token punctuation">(</span>trainset<span class="token punctuation">,</span> batch_size<span class="token operator">=</span><span class="token number">4</span><span class="token punctuation">,</span> shuffle<span class="token operator">=</span><span class="token boolean">True</span><span class="token punctuation">,</span> num_workers<span class="token operator">=</span><span class="token number">2</span><span class="token punctuation">)</span>

<span class="token keyword">class</span> <span class="token class-name">Net</span><span class="token punctuation">(</span>nn<span class="token punctuation">.</span>Module<span class="token punctuation">)</span><span class="token punctuation">:</span>
    <span class="token keyword">def</span> <span class="token function">__init__</span><span class="token punctuation">(</span>self<span class="token punctuation">)</span><span class="token punctuation">:</span>
        <span class="token builtin">super</span><span class="token punctuation">(</span>Net<span class="token punctuation">,</span> self<span class="token punctuation">)</span><span class="token punctuation">.</span>__init__<span class="token punctuation">(</span><span class="token punctuation">)</span>
        self<span class="token punctuation">.</span>conv1 <span class="token operator">=</span> nn<span class="token punctuation">.</span>Conv2d<span class="token punctuation">(</span><span class="token number">1</span><span class="token punctuation">,</span> <span class="token number">10</span><span class="token punctuation">,</span> kernel_size<span class="token operator">=</span><span class="token number">5</span><span class="token punctuation">)</span>
        self<span class="token punctuation">.</span>conv2 <span class="token operator">=</span> nn<span class="token punctuation">.</span>Conv2d<span class="token punctuation">(</span><span class="token number">10</span><span class="token punctuation">,</span> <span class="token number">20</span><span class="token punctuation">,</span> kernel_size<span class="token operator">=</span><span class="token number">5</span><span class="token punctuation">)</span>
        self<span class="token punctuation">.</span>conv2_drop <span class="token operator">=</span> nn<span class="token punctuation">.</span>Dropout2d<span class="token punctuation">(</span><span class="token punctuation">)</span>
        self<span class="token punctuation">.</span>fc1 <span class="token operator">=</span> nn<span class="token punctuation">.</span>Linear<span class="token punctuation">(</span><span class="token number">320</span><span class="token punctuation">,</span> <span class="token number">50</span><span class="token punctuation">)</span>
        self<span class="token punctuation">.</span>fc2 <span class="token operator">=</span> nn<span class="token punctuation">.</span>Linear<span class="token punctuation">(</span><span class="token number">50</span><span class="token punctuation">,</span> <span class="token number">10</span><span class="token punctuation">)</span>

    <span class="token keyword">def</span> <span class="token function">forward</span><span class="token punctuation">(</span>self<span class="token punctuation">,</span> x<span class="token punctuation">)</span><span class="token punctuation">:</span>
        x <span class="token operator">=</span> F<span class="token punctuation">.</span>relu<span class="token punctuation">(</span>F<span class="token punctuation">.</span>max_pool2d<span class="token punctuation">(</span>self<span class="token punctuation">.</span>conv1<span class="token punctuation">(</span>x<span class="token punctuation">)</span><span class="token punctuation">,</span> <span class="token number">2</span><span class="token punctuation">)</span><span class="token punctuation">)</span>
        x <span class="token operator">=</span> F<span class="token punctuation">.</span>relu<span class="token punctuation">(</span>F<span class="token punctuation">.</span>max_pool2d<span class="token punctuation">(</span>self<span class="token punctuation">.</span>conv2_drop<span class="token punctuation">(</span>self<span class="token punctuation">.</span>conv2<span class="token punctuation">(</span>x<span class="token punctuation">)</span><span class="token punctuation">)</span><span class="token punctuation">,</span> <span class="token number">2</span><span class="token punctuation">)</span><span class="token punctuation">)</span>
        x <span class="token operator">=</span> x<span class="token punctuation">.</span>view<span class="token punctuation">(</span><span class="token operator">-</span><span class="token number">1</span><span class="token punctuation">,</span> <span class="token number">320</span><span class="token punctuation">)</span>
        x <span class="token operator">=</span> F<span class="token punctuation">.</span>relu<span class="token punctuation">(</span>self<span class="token punctuation">.</span>fc1<span class="token punctuation">(</span>x<span class="token punctuation">)</span><span class="token punctuation">)</span>
        x <span class="token operator">=</span> F<span class="token punctuation">.</span>dropout<span class="token punctuation">(</span>x<span class="token punctuation">,</span> training<span class="token operator">=</span>self<span class="token punctuation">.</span>training<span class="token punctuation">)</span>
        x <span class="token operator">=</span> self<span class="token punctuation">.</span>fc2<span class="token punctuation">(</span>x<span class="token punctuation">)</span>
        <span class="token keyword">return</span> F<span class="token punctuation">.</span>log_softmax<span class="token punctuation">(</span>x<span class="token punctuation">)</span>

net <span class="token operator">=</span> Net<span class="token punctuation">(</span><span class="token punctuation">)</span>

criterion <span class="token operator">=</span> nn<span class="token punctuation">.</span>CrossEntropyLoss<span class="token punctuation">(</span><span class="token punctuation">)</span>
optimizer <span class="token operator">=</span> optim<span class="token punctuation">.</span>SGD<span class="token punctuation">(</span>net<span class="token punctuation">.</span>parameters<span class="token punctuation">(</span><span class="token punctuation">)</span><span class="token punctuation">,</span> lr<span class="token operator">=</span><span class="token number">0.001</span><span class="token punctuation">,</span> momentum<span class="token operator">=</span><span class="token number">0.9</span><span class="token punctuation">)</span>

<span class="token keyword">for</span> epoch <span class="token keyword">in</span> <span class="token builtin">range</span><span class="token punctuation">(</span><span class="token number">2</span><span class="token punctuation">)</span><span class="token punctuation">:</span>
    running_loss <span class="token operator">=</span> <span class="token number">0.0</span>
    <span class="token keyword">for</span> i<span class="token punctuation">,</span> data <span class="token keyword">in</span> <span class="token builtin">enumerate</span><span class="token punctuation">(</span>trainloader<span class="token punctuation">,</span> <span class="token number">0</span><span class="token punctuation">)</span><span class="token punctuation">:</span>
        <span class="token comment"># NOTE: Data loaded from disk into RAM</span>
        inputs<span class="token punctuation">,</span> labels <span class="token operator">=</span> data

        optimizer<span class="token punctuation">.</span>zero_grad<span class="token punctuation">(</span><span class="token punctuation">)</span>

        <span class="token comment"># NOTE: CPU computes forward pass, backward pass, update weights</span>
        outputs <span class="token operator">=</span> net<span class="token punctuation">(</span>inputs<span class="token punctuation">)</span>
        loss <span class="token operator">=</span> criterion<span class="token punctuation">(</span>outputs<span class="token punctuation">,</span> labels<span class="token punctuation">)</span>
        loss<span class="token punctuation">.</span>backward<span class="token punctuation">(</span><span class="token punctuation">)</span>
        optimizer<span class="token punctuation">.</span>step<span class="token punctuation">(</span><span class="token punctuation">)</span>

        running_loss <span class="token operator">+=</span> loss<span class="token punctuation">.</span>item<span class="token punctuation">(</span><span class="token punctuation">)</span>
        <span class="token keyword">if</span> i <span class="token operator">%</span> <span class="token number">2000</span> <span class="token operator">==</span> <span class="token number">1999</span><span class="token punctuation">:</span>
            <span class="token keyword">print</span><span class="token punctuation">(</span><span class="token string">'[%d, %5d] loss: %.3f'</span> <span class="token operator">%</span> <span class="token punctuation">(</span>epoch <span class="token operator">+</span> <span class="token number">1</span><span class="token punctuation">,</span> i <span class="token operator">+</span> <span class="token number">1</span><span class="token punctuation">,</span> running_loss <span class="token operator">/</span> <span class="token number">2000</span><span class="token punctuation">)</span><span class="token punctuation">)</span>
            running_loss <span class="token operator">=</span> <span class="token number">0.0</span>
</code></pre>
<p>Several things verify that our model works correctly:</p>
<ol>
<li>There were no interpreter errors when we run the code (i.e. no crashes)</li>
<li>The loss decreases</li>
<li>The accuracy on the validation increases</li>
</ol>
<p>This example is representative of how tutorials and most universities teach deep learning, but it is not useful in practice. Let's dig in on several of the assumptions that make this example so elegant. First, the MNIST dataset is <strong>tiny</strong>, only 9.9 MB,  and <strong>static</strong>. Why do we focus on this? Because in practice, the datasets we want to use to power insights or new features are <strong>massive</strong> and <strong>dynamic</strong> (constantly changing). It is easy in this example to download MNIST to our local device to train the model - it only takes a few seconds to make an HTTP request to the server with the tarball of the MNIST dataset and download the whole dataset. And once the dataset is in memory, the time cost of loading small batches of the dataset into our model on the CPU is <strong>effectively optimal</strong>. MNIST is just 28x28 pixel images - it takes almost no time to load a small batch of 28x28 images, and those images take up almost no RAM. You can't beat the speed of just loading data from disk onto CPUs without fancy I/O optimizations which we'll cover in a later section. If we were lazy and inefficient, we could always re-download the MNIST dataset from the server every time we trained the model and <strong>it would still work</strong>. We would only add a few seconds to each training run, and the amount of memory we consume is only 9.9 MB.</p>
<p>Also observe that all of the operations in the function above.. TODO</p>
<h2>Distributed Deep Learning</h2>
<p>Now we that we've shown the example above, we can <em>throw away almost every assumption we just made</em>. In real life, DL / AI / ML is not magic. All we are doing is using backpropagation to compute gradient updates for an optimization function. The <em>magic</em> is cleverly selecting a model that maps our input data to our objective, sometimes being clever with <em>how</em> we train the model(s), and training on as much data we have available. For all of the literature on "efficient training" and network architectures that have stronger capacity to <em>generalize</em>, in practice we want to train the model using <strong>as much data as humanly possible</strong>. Practically any clever DL trick we can think of can be ignored if we have more training data. So if you're a DL practitioner in an organization, and the goal is to get a model you can actually use (in business or research), your main goal is to acquire as much data as possible. If you're successful and can acquire a massive dataset (&gt;10TB), you will be able to train something useful, but your dataset <strong>does not fit in memory</strong>. Furthermore, your dataset is probably generated via an internal ETL pipeline, or by interns tasked with getting you more data, which means it <strong>changes over time</strong>.</p>
<p>It will be useful to understand the role of distributed communications in DL before digging into distributed deep training. To speed up model training on a large dataset using multiple GPUs, we turn to "<em>data parallel training</em>". The plain english explanation is we can speed up model training by sending different chunks (i.e. "<em>batches</em>") of the very large dataset to each GPU, have each GPU compute weight updates separately, and then add up those weight updates (i.e. "<em>gradients</em>") to produce one global weight update at each time step. The step where we add up the weight updates is known as <code>all_reduce()</code>, and we'll cover it in more detail in the MPI section. The copy of the model on each GPU will have the same weights after each iteration of backpropagation, but we've now consumed <em># gpu's</em> times the data in a single iteration. Throughout this process the GPU's need to be in constant communication, and so a distributed communications protocol is required.</p>
<img alt="Distributed Communication." loading="lazy" width="770" height="401" decoding="async" data-nimg="1" style="color:transparent" src="/_next/static/media/DistributedBackend.e44c67b0.svg?dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW">
<h3>Hardware</h3>
<p>The bottom layer of our distributed training stack is hardware. Our network has CPUs, GPUs, and in the future perhaps TPUs and FPGAs. For modern deep learning, the only GPUs anyone uses are sold by NVIDIA - they hold a monopoly on the market. Why? How? The short story: NVIDIA owns CUDA. CUDA makes it easy to write code that can take advantage of GPU parallelization. Deep learning, and especially computer vision based deep learning, is <strong>a lot of convolutions</strong>. Convolutions are "shift, add, sum" - very easy math to parallelize, <strong>extremely fast operation when parallelized</strong>. All the major numerical computation libraries built in CUDA support to use GPUs since it was first and easiest. Only NVIDIA GPUs support CUDA, so by default you can now only use NVIDIA GPUs for deep learning.</p>
<blockquote>
<p>CUDA is a parallel computing platform and programming model that makes using a GPU for general purpose computing simple and elegant. The developer still programs in the familiar C, C++, Fortran, or an ever expanding list of supported languages, and incorporates extensions of these languages in the form of a few basic keywords.... These keywords let the developer express massive amounts of parallelism and direct the compiler to the portion of the application that maps to the GPU.</p>
</blockquote>
<p>In 2016 Google announced the Tensor Processing Unit (TPU), an application specific integrated circuit (ASIC) specifically built for deep learning. Without going into hardware detail, TPUs are significantly faster than GPUs for deep learning. GPUs are multipurpose - they are graphics processing units, literally designed to handle graphics workloads. It just so happens that GPUs are effective for deep learning, but they are not power or memory efficient. Unlike GPUs, TPUs leverage <code>systolic arrays</code> to handle large multiplications and additions with memory efficiency <sup><a href="#user-content-fn-1" id="user-content-fnref-1" data-footnote-ref="true" aria-describedby="footnote-label">1</a></sup>. Google initially used its TPUs internally to handle its massive DL workloads, but now offers TPU enabled training and inference as a cloud offering and for sale.</p>
<h3>Network Switching</h3>
<h3>Message Passing Interface (MPI)</h3>
<p>Message Passing Interface (MPI) surfaced out of a supercomputing computing community working group in the 1990's. From the MPI 4.0 specification<sup><a href="#user-content-fn-2" id="user-content-fnref-2" data-footnote-ref="true" aria-describedby="footnote-label">2</a></sup>:</p>
<blockquote>
<p>MPI (Message-Passing Interface) is a message-passing library interface specification...
MPI addresses primarily the message-passing parallel programming model, in which data is moved from the address space of one process to
that of another process through cooperative operations on each process. Extensions to the
“classical” message-passing model are provided in collective operations, remote-memory
access operations, dynamic process creation, and parallel I/O</p>
</blockquote>
<p>MPI specifies <em>primitives</em> (i.e. building blocks) that can be used to compose complex distributed communications patterns. MPI specifies two main types of primitives:</p>
<ol>
<li><em>point-to-point</em> : one process communicates with another process (1:1)</li>
<li><em>collective</em> : group of processes communicates within the group (many:many)</li>
</ol>
<p>These primitives are obviously useful for distributed deep learning. Each GPU process has weight gradients we need add up? Use the <code>all_reduce()</code> primitive. Each GPU process <code>P_i</code> comes up with a tensor <code>x_i</code> that needs to be shared to all other GPU processes? Use the <code>all_gather()</code> primitive, etc. Below is a helpful graphic directly from the MPI 4.0 standard to cement the concept.</p>
<img alt="Local Global rank" loading="lazy" width="960" height="720" decoding="async" data-nimg="1" style="color:transparent" src="/_next/static/media/LocalGlobalRank.97d7d984.svg?dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW">
<img alt="Collective Communication" loading="lazy" width="478" height="700" decoding="async" data-nimg="1" style="color:transparent" srcset="/_next/image?url=%2F_next%2Fstatic%2Fmedia%2FMPICollectiveCommunication.b26b15ff.png&amp;w=640&amp;q=75&amp;dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW 1x, /_next/image?url=%2F_next%2Fstatic%2Fmedia%2FMPICollectiveCommunication.b26b15ff.png&amp;w=1080&amp;q=75&amp;dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW 2x" src="/_next/image?url=%2F_next%2Fstatic%2Fmedia%2FMPICollectiveCommunication.b26b15ff.png&amp;w=1080&amp;q=75&amp;dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW">
<h3>Distributed Training Frameworks</h3>
<p>Horovod: <a href="https://youtu.be/SphfeTl70MI">https://youtu.be/SphfeTl70MI</a></p>
<p>Comparison of NVLink vs alternatives: <a href="https://arxiv.org/pdf/1903.04611.pdf">https://arxiv.org/pdf/1903.04611.pdf</a></p>
<h2>Single GPU Training</h2>
<p>This is the same scenario as "CPU Training on Machine", but we now have 1 GPU in addition to our CPUs.</p>
<img alt="Single GPU" loading="lazy" width="1369" height="727" decoding="async" data-nimg="1" style="color:transparent" src="/_next/static/media/SingleGPU.da7a5148.svg?dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW">
<p><a href="https://pytorch.org/docs/stable/data.html#torch.utils.data.DataLoader">https://pytorch.org/docs/stable/data.html#torch.utils.data.DataLoader</a></p>
<h2>Multi GPU, Single Node Training</h2>
<p>This is the same scenario as "Single GPU Training", but we now have &gt;1 GPUs on our node. Most deep learning practitioners in academia, small industry research groups, and hobbyists operate under this scenario. Simply put, you take as many GPU's as you can afford and put them on a single physical machine. This avoids the need to consider inter-machine networking, and in reality, this still works for most SOTA research.</p>
<p>As seen in the figure below, the deep learning framework creates a process to manage each GPU's training. The GPU's each receive a different mini-batch of the data for the forward pass, compute gradients, and then proceed to <code>all_reduce</code>. Now we have two main ways to perform the <code>all_reduce</code>.</p>
<p>The first way is to designate one of the processes the <code>master</code> process, usually the process managing GPU 0, and have each GPU process send gradients to the <code>master</code> process. The master process accumulates all gradients and sends the <em>reduced</em> gradients back to each GPU process. Now each GPU has the same gradients to update weights. This is particularly effective for this scenario because the GPUs are all <strong>on the same machine</strong>; they are connected by a dedicated IO bus that supports extremely fast inter-GPU communication. In PyTorch this the <code>DataParallel</code>[^3] model</p>
<p>TODO: Does the second way exist??? What are the tradeoffs if so.</p>
<img alt="Node" loading="lazy" width="1412" height="900" decoding="async" data-nimg="1" style="color:transparent" src="/_next/static/media/MultiGPUMultiNode.080a402c.svg?dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW">
<p><a href="https://pytorch.org/tutorials/intermediate/dist_tuto.html">https://pytorch.org/tutorials/intermediate/dist_tuto.html</a>
<a href="https://pytorch.org/docs/stable/distributed.html">https://pytorch.org/docs/stable/distributed.html</a></p>
<h2>Multi GPU, Multi Node Training</h2>
<p>All processes in the process group<sup><a href="#user-content-fn-4" id="user-content-fnref-4" data-footnote-ref="true" aria-describedby="footnote-label">3</a></sup> use ring communication</p>
<img alt="Multi GPU" loading="lazy" width="1392" height="725" decoding="async" data-nimg="1" style="color:transparent" src="/_next/static/media/MultiGPU.7adc1f7f.svg?dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW">
<h2>Multi GPU, Multi-Node Training Cloud</h2>
<p>Truly large-scale models are trained in the multi-gpu, multi-node setting. Examples include GPT-3, DeepMind's DOTA model, and Google's DL driven recommendation algorithms. Recall that a 'node' is typically a physical machine with 8 GPU's. When practitioners say a model was trained on 1024 GPU's, that precisely means the model was trained on 1024 / 8 = 128 physical nodes, each with 8 GPU's.</p>
<p>Let's take a moment to think of the broad challenges here before getting into specifics. We're now training a model across 128 separate physical machines. This introduces problems related to communication, i.e. that the physical nodes will have to communicate with each other over some network, and problems of speed. It takes time to load each batch of data from the data store to each physical machine. What if the data size is large (videos, raw audio), and what if DB reads are slow?</p>
<p>6a. Distributed Training Frameworks (<em>Training Backends</em>)
a. MPI
b. DDP (NCCL)
c. Horovod
<a href="https://github.com/horovod/horovod">https://github.com/horovod/horovod</a>
d. BytePS
e. Parameter Servers</p>
<p><a href="https://pytorch.org/tutorials/intermediate/rpc_param_server_tutorial.html">https://pytorch.org/tutorials/intermediate/rpc_param_server_tutorial.html</a></p>
<h2>Cloud Workflows for Deep Learning</h2>
<h2>Deep Learning Frameworks</h2>
<h2>Higher Level Frameworks</h2>
<p>Keras, Ignite, Lightning</p>
<h2>Data Ingestion Patterns</h2>
<h2>Frameworks*</h2>
<p>Let's talk about frameworks for a second. This section has an asterisk because I want to make it clear these are opinions, not facts. Frameworks provide higher level patterns / abstractions that structure your usage of a tool and <em>in theory</em> enable you to do more complex tasks with less effort. I say in theory because at worst, the frameworks themselves have a steep learning curve and don't enable you to do anything the underlying tool couldn't do.</p>
<h3>PyTorch Ignite with PyTorch</h3>
<ul>
<li>Elegant <code>engine</code> construct to manage logging, hooks, training execution</li>
<li>Clean interface over PyTorch distributed module</li>
<li>Narrow set of excellent utilities for logging, checkpoint management, etc.</li>
<li>Poor support for custom code - rewriting open source code to operate with an ignite engine is difficult</li>
<li>Low-level, sits one level above PyTorch so not <em>too</em> abstract</li>
<li>Small but active developer community</li>
</ul>
<h3>PyTorch Lightning with PyTorch</h3>
<ul>
<li>Strong abstractions for new developers</li>
<li>Excellent support for multi GPU training</li>
<li>Integrations with upcoming deep learning backends</li>
<li>Large developer community</li>
<li>The guy who started PyTorch Lightning is shameless, he goes for personal attention at every opportunity, and they misrepresent other's work as the Lighning's work. -10.</li>
</ul>
<h3>Keras with TensorFlow</h3>
<ul>
<li>Excellent abstractions over TensorFlow for new developers</li>
<li>Such widespread usage that Google moved to provide first-class support</li>
<li>TensorFlow adopted many core ideas</li>
<li>No longer needed to advanced TensorFlow</li>
<li>Doesn't play nice with the entire Google Cloud DL ecosystem</li>
</ul>
<h2>Reading List (Links)</h2>
<ul>
<li>Andrej Karpathy's Recipe for Training Neural Networks (2019)</li>
<li><a href="https://lambdalabs.com/blog/introduction-multi-gpu-multi-node-distributed-training-nccl-2-0/">https://lambdalabs.com/blog/introduction-multi-gpu-multi-node-distributed-training-nccl-2-0/</a></li>
</ul>
<section data-footnotes="true" class="footnotes"><h2 class="sr-only" id="footnote-label">Footnotes</h2>
<ol>
<li id="user-content-fn-1">
<p><a href="https://cloud.google.com/blog/products/ai-machine-learning/what-makes-tpus-fine-tuned-for-deep-learning">https://cloud.google.com/blog/products/ai-machine-learning/what-makes-tpus-fine-tuned-for-deep-learning</a> <a href="#user-content-fnref-1" data-footnote-backref="" aria-label="Back to reference 1" class="data-footnote-backref">↩</a></p>
</li>
<li id="user-content-fn-2">
<p><a href="https://www.mpi-forum.org/docs/mpi-4.0/mpi40-report.pdf">https://www.mpi-forum.org/docs/mpi-4.0/mpi40-report.pdf</a> <a href="#user-content-fnref-2" data-footnote-backref="" aria-label="Back to reference 2" class="data-footnote-backref">↩</a></p>
</li>
<li id="user-content-fn-4">
<p>https://NEED_THE_LINK_FOR_DATA_PARALLEL.com <a href="#user-content-fnref-4" data-footnote-backref="" aria-label="Back to reference 3" class="data-footnote-backref">↩</a></p>
</li>
</ol>
</section>]]></content:encoded>
            <author>gmgprivacy@proton.me (Matt Goodman)</author>
        </item>
        <item>
            <title><![CDATA[Mega-guide to Nvidia CUDA, CUDA Toolkit, CUDA Driver for an Ubuntu 20.04 Nvidia GPU PC Build]]></title>
            <link>https://goodmattg.xyz/articles/Mega-guide-CUDA</link>
            <guid>https://goodmattg.xyz/articles/Mega-guide-CUDA</guid>
            <pubDate>Sat, 25 Jun 2022 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>Setting up an Nvidia GPU enabled system is the ante for deep learning research at home. Without a dedicated GPU, users are forced to either use Google Colab, which has great performance for a free service but is a notebook environment, or a cloud ML service that costs money (~$0.9 GPU/hour on AWS). After playing with those options for awhile, I decided to build my own system. In doing so, I learned how confusing the process can be, especially if it’s your first PC build or first build with a dedicated GPU system.</p>
<p>So before getting into this a note of praise - you aren’t dumb if you don’t fully understand CUDA drivers. I’m a deep learning practitioner with software experience, and this was painful for me.</p>
<p>First some terminology:</p>
<ul>
<li><strong>Nvidia X.Y Toolkit</strong>: what tutorials mean when they say “install CUDA X.Y”. The core versions of CUDA are 10.xx (past) and 11.yy (future)</li>
<li><strong>Nvidia Driver</strong>: the driver controlling defining GPU’s interface with your OS / Architecture.</li>
<li><code>nvcc</code>: the CUDA compiler driver</li>
</ul>
<p>Nvidia Software Ecosystem at-a-glance.</p>
<img alt="Nvidia Software Ecosystem at-a-glance." loading="lazy" width="950" height="598" decoding="async" data-nimg="1" style="color:transparent" srcset="/_next/image?url=%2F_next%2Fstatic%2Fmedia%2FCUDA.a893e8cc.png&amp;w=1080&amp;q=75&amp;dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW 1x, /_next/image?url=%2F_next%2Fstatic%2Fmedia%2FCUDA.a893e8cc.png&amp;w=1920&amp;q=75&amp;dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW 2x" src="/_next/image?url=%2F_next%2Fstatic%2Fmedia%2FCUDA.a893e8cc.png&amp;w=1920&amp;q=75&amp;dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW">
<p>Toolkit documentation:</p>
<p><a href="https://docs.nvidia.com/cuda/cuda-toolkit-release-notes/index.html">Release Notes :: CUDA Toolkit Documentation</a></p>
<p>Also useful is the driver documentation. You can find it for your specific driver version at:</p>
<p><a href="https://download.nvidia.com/XFree86/Linux-x86_64/515.43.04/README/">NVIDIA Accelerated Linux Graphics Driver README and Installation Guide</a></p>
<h1>The Decision: picking your CUDA installation</h1>
<p>If you read the release notes above, you’ll observe multiple installation methods for the CUDA driver / toolkit pair on linux. It can be deeply confusing. Do I go with the runfile or Debian installer? Should I just go with the newest versions of the CUDA toolkit and CUDA driver?</p>
<p>Nvidia makes it clear that there is no correct answer here. While there are some constraints (described below), just do whatever is right for your system.</p>
<h3>CUDA Compatibility</h3>
<p>Not every version of CUDA driver supports every version of the CUDA Toolkit. Before making a choice, refer to CUDA compatibility matrix, copied here for reference.</p>
<img loading="lazy" width="1148" height="568" decoding="async" data-nimg="1" style="color:transparent" srcset="/_next/image?url=%2F_next%2Fstatic%2Fmedia%2FCUDA_MINOR_VERSION_COMPAT.d83146a7.png&amp;w=1200&amp;q=75&amp;dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW 1x, /_next/image?url=%2F_next%2Fstatic%2Fmedia%2FCUDA_MINOR_VERSION_COMPAT.d83146a7.png&amp;w=3840&amp;q=75&amp;dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW 2x" src="/_next/image?url=%2F_next%2Fstatic%2Fmedia%2FCUDA_MINOR_VERSION_COMPAT.d83146a7.png&amp;w=3840&amp;q=75&amp;dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW">
<p><a href="https://docs.nvidia.com/deploy/cuda-compatibility/index.html">NVIDIA Driver Documentation</a></p>
<p><strong>But wait!</strong> This is not the actual table you should refer to - this table just says that for CUDA Toolkit 11.x minor version compatibility (i.e. to even be able to support minor version “x”), you need at least driver ≥450.80.02. The actual table we will refer to is below.</p>
<img loading="lazy" width="1216" height="1520" decoding="async" data-nimg="1" style="color:transparent" srcset="/_next/image?url=%2F_next%2Fstatic%2Fmedia%2FCUDA_VERSION_COMPAT.e6479778.png&amp;w=1920&amp;q=75&amp;dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW 1x, /_next/image?url=%2F_next%2Fstatic%2Fmedia%2FCUDA_VERSION_COMPAT.e6479778.png&amp;w=3840&amp;q=75&amp;dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW 2x" src="/_next/image?url=%2F_next%2Fstatic%2Fmedia%2FCUDA_VERSION_COMPAT.e6479778.png&amp;w=3840&amp;q=75&amp;dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW">
<p>This is the reference table we care about - it shows the minimum driver required for Linux 64 bit x86 architecture to support a given CUDA Toolkit version. Let’s say we decide on 11.2.0 GA for the global installation of CUDA Toolkit on our system. We will need at least CUDA driver 460.27.03 on our system. That’s it, we’ve solved it 🤝. We’re free to use a higher version of the CUDA driver, since CUDA drivers are always designed to be backwards compatible, but sticking with a lower (stable) version of driver is fine too.</p>
<h2>Special Case: Forward Compatibility</h2>
<p>Let’s quickly discuss a special use-case. The vast majority of users will go with a stable pair of CUDA toolkit and CUDA driver, but what about the case where we’re fixed to an old driver but want to use a newer version of CUDA.</p>
<p>As an example, let’s say we are fixed to CUDA 418.40.04, but need our system to use CUDA 11.2. The compatibility table says we are out of luck, since the minimum required version is 460.27.03 😩. Forward compatibility packages to the rescue! Referring to the table below, we see forward compability for the 418.40.04 driver is available for CUDA 10.2-11.6 🎉.</p>
<img loading="lazy" width="1904" height="684" decoding="async" data-nimg="1" style="color:transparent" srcset="/_next/image?url=%2F_next%2Fstatic%2Fmedia%2FFORWARD_COMPATIBILITY.a541ce87.png&amp;w=1920&amp;q=75&amp;dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW 1x, /_next/image?url=%2F_next%2Fstatic%2Fmedia%2FFORWARD_COMPATIBILITY.a541ce87.png&amp;w=3840&amp;q=75&amp;dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW 2x" src="/_next/image?url=%2F_next%2Fstatic%2Fmedia%2FFORWARD_COMPATIBILITY.a541ce87.png&amp;w=3840&amp;q=75&amp;dpl=dpl_4LqCHZBSNhvRUEKFcppYv1Xij5AW">
<p>All we have to do is install the forward compatiblity package onto our system. On Ubuntu, prefer installing the compatibility package from network repositories. In highly specific use cases, can install the compatibility package manually from a runfile (.run) - but for most users just use <code>apt-get</code> to install.</p>
<pre class="language-bash"><code class="language-bash"><span class="token function">sudo</span> <span class="token function">apt-get</span> <span class="token function">install</span> -y cuda-compat-11-2
</code></pre>
<p>The package will install to <code>/usr/local/cuda/compat</code> . Now include the installed CUDA compability files. This code below also runs the CUDA Device Query utility which “enumerates the properties of the CUDA devices present in the system”.****</p>
<pre class="language-bash"><code class="language-bash"><span class="token assign-left variable">LD_LIBRARY_PATH</span><span class="token operator">=</span>/usr/local/cuda/compat:<span class="token variable">$LD_LIBRARY_PATH</span> samples/bin/x86_64/linux/release/deviceQuery
</code></pre>
<h2>The Installation</h2>
<p>Quoting from the Nvidia <a href="https://docs.nvidia.com/cuda/cuda-quick-start-guide/index.html#ubuntu-x86_64">docs</a>:</p>
<blockquote>
<p>When installing CUDA on Ubuntu, you can choose between the Runfile Installer and the Debian Installer. The Runfile Installer is only available as a Local Installer. The Debian Installer is available as both a Local Installer and a Network Installer. The Network Installer allows you to download only the files you need. The Local Installer is a stand-alone installer with a large initial download. In the case of the Debian installers, the instructions for the Local and Network variants are the same.</p>
</blockquote>
<p>As a heuristic, if you don’t know why you would choose one installation method over the other, <strong>choose the Debian Installer!</strong></p>
<p>I’m not going to copy steps from the docs, but to summarize:</p>
<ol>
<li>Go through <a href="https://docs.nvidia.com/cuda/cuda-installation-guide-linux/index.html#pre-installation-actions">pre-installation</a> actions<!-- -->
<ol>
<li>If going through a “local” installation (.deb or .runfile install), instead of network installation (.deb install only), download the installer to your system from the <a href="https://developer.nvidia.com/cuda-toolkit-archive">downloads</a> page for the specific CUDA toolkit you want to install</li>
</ol>
</li>
<li>Go through <a href="https://docs.nvidia.com/cuda/cuda-installation-guide-linux/index.html#ubuntu-installation">installation</a> action</li>
<li>Go through <a href="https://docs.nvidia.com/cuda/cuda-installation-guide-linux/index.html#post-installation-actions">post-installation</a> actions</li>
</ol>
<h1>Hard Reset: clearing your system of installed CUDA packages</h1>
<p>“Removing CUDA” is a popular topic online. Now that we understand CUDA is actually a collection of packages (Toolkit, Driver, forward-compat, etc.) we have a better idea of what it takes to uninstall CUDA. First, from the Nvidia Docs we can uninstall a CUDA <code>deb</code> installation on Ubuntu using:</p>
<pre class="language-bash"><code class="language-bash"><span class="token function">sudo</span> <span class="token function">apt-get</span> --purge remove cuda*
</code></pre>
<p>Remember that using <code>deb</code> installs files to <code>/usr/local</code>, so after this command <code>ls /usr/local/ | grep cuda</code> should show no matched files.</p>
<p>If we installed using a runfile instead of the deb installer, we need to uninstall the Driver and Toolkit with separate commands:</p>
<pre class="language-bash"><code class="language-bash"><span class="token function">sudo</span> /usr/local/cuda-X.Y/bin/cuda-uninstaller <span class="token comment"># Uninstall the Toolkit</span>
<span class="token function">sudo</span> /usr/bin/nvidia-uninstall <span class="token comment"># Uninstall the Driver</span>
</code></pre>
<h1>Anaconda: the easiest way to manage CUDA versions across projects</h1>
<p>The easiest way to manage CUDA for different DL projects is manage your environments with Anaconda.</p>
<p>Some highlights:</p>
<ul>
<li>Anaconda can manage multiple CUDA toolkit versions separately from your system’s base CUDA toolkit version</li>
<li>Anaconda takes care of the overhead of installation and un-installation.</li>
<li>Uninstall is as easy as <code>conda remove cuda</code></li>
</ul>
<h1>Attempt #1</h1>
<p>I had almost no idea what I was doing for this attempt. My only requirement was CUDA 11.7 because I wanted to run Jax from the base CUDA system install and Jax requires 11.7. Given this tutorial, I now know this is not smart. You can install a system-wide version of Toolkit + Driver that is stable to the system hardware (i.e. support suspend-to-RAM) and use conda to manage a higher CUDA version that supports Jax. Installed CUDA driver 515 and CUDA toolkit 11.7 via the Runfile installation method. <code>deviceQuery</code> sample returns expected output and <code>nvidia-smi</code> command also shows expected output. However, I consider this installation attempt a failure. Most important feature for my system is <code>suspend-to-RAM</code> - i.e. the ability to take GPU state at a given time, move it to RAM, and put the system into “suspend” mode without fully turning off the power. This is the most important feature to modern computer usage - i.e. shutting your laptop lid, opening back up and having all the applications you had open... still open. Having to do a full power cycle after every computer use significantly hampers usability 🤦🏼‍♂️. I’ll have to purge everything and try again.</p>
<h1>Attempt #2 (ongoing)</h1>
<p>Pre-notes: I’m going to try purging all CUDA from the entire system first. I’ve set in the BIOS that the system uses integrated graphics on the mobo, not the GPU, for graphics, so the system will boot fine without the GPU drivers installed. I have a feeling using the newest Nvidia driver (515) is the problem - going to try reverting to 470 since it comes with the Debian installer and is the base version installed on Ubuntu by default. I’m also hoping using the Debian installer fixes the “suspend-to-RAM” issue. If you read the release notes for the drivers (470 vs. 515) they handle power management differently 🤔. I’m sure 515 has more features + efficiency, but I want things to work.</p>
<h1>Summary</h1>
<p>Installing CUDA on your system is non-trivial. In general, the Nvidia docs are thorough and <strong>do</strong> contain all of the information you might need, but feel inscrutable to a first-time user. It would probably be helpful for Nvidia to put together more quick installation guides for common configurations. This guide is by no means complete (or even correct)! If I’m missing anything, please let me know.</p>]]></content:encoded>
            <author>gmgprivacy@proton.me (Matt Goodman)</author>
        </item>
        <item>
            <title><![CDATA[Mythos and DeFi]]></title>
            <link>https://goodmattg.xyz/articles/Mythos-and-DeFi</link>
            <guid>https://goodmattg.xyz/articles/Mythos-and-DeFi</guid>
            <pubDate>Fri, 24 Apr 2026 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>If you’re like me and you ingest media compulsively it feels like we’re standing at a tipping point with release of <a href="https://www.anthropic.com/glasswing">Mythos</a>. This is either the greatest marketing strategy ever (“so powerful we can’t let you use it”) or we’ve made another leap on our acceleration to AGI (see my <a href="/articles/Ashenbrenner-Review">other article</a>). My guess is with the release of <a href="https://openai.com/index/introducing-gpt-5-5/">GTP 5.5</a> with comparable <a href="https://openai.com/index/introducing-gpt-5-5/">Terminal-Bench 2.0</a> results at 82.7%, it’s all hype in the sense that no, this is not the model that immediately self-accelerates to destroy us all, but that yes, every 5-7% improvement in Terminal -Bench / Expert-SWE performance affirms the utility of these models.</p>
<p>What I’m amused by is how the first thing the media jumps to is zero days in our critical software infrastructure that is open source (the <a href="https://github.com/torvalds/linux">linux kernel</a>, etc), and they’re right, that is scary. Because if all it takes is a security researcher saying “find me a zero day plz”, is 100 days really enough to patch everything? I truly doubt it, because the first thing red team will do when the model hits open access is ask:</p>
<p>“For each of the next 10000k non-famous, but still critical open source projects, find me zero days please”.</p>
<p>And if the <a href="https://www.anthropic.com/glasswing">Mythos</a> claims are real, that WILL work with certainty. And if it doesn’t on the first pass, we’ll just keep doing it again and again and again telling the model to be more exhaustive. Because this is 100 days of wall clock time but Anthropic is obviously limiting the compute over those 100 days due to costs and resource constraints, and the red team will have infinite time after the model release.</p>
<p>As little as I care about using Defi myself, I’ll be watching the markets  after this release. Think about it - what other markets immediately tie economic value to Mythos token. If Mythos tokens really are this magic vulnerability finding machine, then we’ll immediately see those tokens be used to extract economic value from the anonymous internet UNTIL the token values surpass the economic value to be extracted OR the TVL locked in these Defi protocols and contracts decreases. As I’m writing this the <a href="https://coinmarketcap.com/view/defi/">total Defi market cap</a> is 77.21 BUSD - so if those protocols start getting breached due to the new omni-model, we’ll see the effects immediately, but I’m curious how quickly that will actually happen. As the low hanging vulnerability fruit is picked, how much token kerosene are we going to burn searching for new vulnerabilities that translate immediately to money?</p>
<div class="not-prose my-8 overflow-hidden rounded border border-zinc-200 bg-white dark:border-zinc-800 dark:bg-zinc-950"><iframe title="TradingView Crypto Total DeFi Market Cap chart" src="https://s.tradingview.com/widgetembed/?symbol=CRYPTOCAP%3ATOTALDEFI&amp;interval=D&amp;range=3M&amp;theme=light&amp;style=1&amp;timezone=Etc%2FUTC&amp;withdateranges=1&amp;hideideas=1&amp;hidesidetoolbar=1&amp;saveimage=0" class="h-[420px] w-full"></iframe></div>
<p><a href="https://www.tradingview.com/symbols/TOTALDEFI/">Source: TradingView CRYPTOCAP:TOTALDEFI</a></p>]]></content:encoded>
            <author>gmgprivacy@proton.me (Matt Goodman)</author>
        </item>
        <item>
            <title><![CDATA[Lessons Learned using MyPy in Production]]></title>
            <link>https://goodmattg.xyz/articles/Safe-Typing-In-Python</link>
            <guid>https://goodmattg.xyz/articles/Safe-Typing-In-Python</guid>
            <pubDate>Wed, 07 Apr 2021 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<p>When we talk about type safety in the context of statically typed languages, we mean that the language builds in a typing checking mechanism into the compiler. No type-check, no compilation.</p>
<p>But in the context of research / algorithms oriented software that has no uptime requirements, the bias coming out of academia is to get the math into code, these days Python, and put it in production. Python is dynamically typed and has no type inference... so the typing is optional. What Mypy users rarely discuss is the cognitive burden of maintaining the type system that is separate from the code. With a lot of tasks on your plate and new features you are waiting to start, sometimes it feels silly to type annotate Python. Is this really necessary?</p>
<pre class="language-python"><code class="language-python"><span class="token keyword">def</span><span class="token punctuation">(</span>a<span class="token punctuation">)</span><span class="token punctuation">:</span>
    <span class="token keyword">print</span><span class="token punctuation">(</span>a<span class="token punctuation">)</span>
    <span class="token keyword">return</span>

<span class="token comment"># With types</span>
<span class="token keyword">def</span><span class="token punctuation">(</span>a<span class="token punctuation">:</span> <span class="token builtin">int</span><span class="token punctuation">)</span> <span class="token operator">-</span><span class="token operator">&gt;</span> <span class="token boolean">None</span><span class="token punctuation">:</span>
    <span class="token keyword">print</span><span class="token punctuation">(</span>a<span class="token punctuation">)</span>
    <span class="token keyword">return</span>
</code></pre>
<p>The situation we experience more often is that the simple and obvious functions get type annotated, but the complex ones that deal with uncommon types, like declaring types for decorators <code>F = TypeVar('F', bound=Callable[..., Any])</code>, go un-typed. And we expect this! Catching subtle bugs via the mypy type-checking mechanism is a long game.</p>
<p>Highly effective teams know that type annotation is one of the best ways to make a codebase <em>stronger</em>. I know that's a vague and probably incorrect word to use in the context of software, but it really feels that way in a Python codebase. Without type annotations, it's the blind leading the blind - you need faith that variable names accurately describe the data that code will operate on. "Faith" is cold-comfort when we start talking SLA's.</p>
<pre class="language-python"><code class="language-python"><span class="token keyword">def</span> <span class="token function">concatenate_str</span><span class="token punctuation">(</span>x<span class="token punctuation">:</span> <span class="token builtin">str</span><span class="token punctuation">,</span> y<span class="token punctuation">:</span> <span class="token builtin">str</span><span class="token punctuation">)</span> <span class="token operator">-</span><span class="token operator">&gt;</span> <span class="token builtin">str</span><span class="token punctuation">:</span>
    <span class="token keyword">return</span> x <span class="token operator">+</span> y

<span class="token comment"># the function name said '_str' but our data is 'int'</span>
concatenate<span class="token punctuation">(</span><span class="token number">1</span><span class="token punctuation">,</span> <span class="token number">2</span><span class="token punctuation">)</span> <span class="token operator">-</span><span class="token operator">&gt;</span> <span class="token number">3</span>
</code></pre>
<p>prevent programmer's from making dumb mistakes in their work. I believe that well-designed code is a force multiplier</p>
<p>The following code will type check, even though the types are wrong. Mypy is a tool for type <strong>annotation</strong>; it's a document of what you want the types to be, and if the types are what you say they are, then the code is type safe.</p>
<h1>Lessons</h1>
<p>So if the investment in type annotations (Mypy) only pays off with a full commitment, the most common scientific libraries are still working on type stubs, and no one has the time or willingness to fully commit, the question is, why even bother?</p>
<h2>1. Mindfulness is enough to prevent common mistakes</h2>
<p>Even if they go completely unused in any formal way, just making developers write types prevents dumb mistakes. For Python, this is often <code>Optional[T]</code> vs <code>[T]</code>. <strong>This is not a joke</strong>. Errors like this have led to downtime and rollbacks for production services.</p>
<blockquote>
<p>"operator '&gt;' not defined for Int and None"</p>
</blockquote>
<pre class="language-python"><code class="language-python"><span class="token comment"># We have a vague understanding that 'a' is an int</span>
__init__<span class="token punctuation">(</span>self<span class="token punctuation">,</span> val<span class="token operator">=</span><span class="token boolean">None</span><span class="token punctuation">)</span><span class="token punctuation">:</span>
    func<span class="token punctuation">(</span>val<span class="token punctuation">)</span>

<span class="token comment"># No types - explosion</span>
<span class="token keyword">def</span> <span class="token function">func</span><span class="token punctuation">(</span>a<span class="token punctuation">)</span><span class="token punctuation">:</span>
    <span class="token keyword">if</span> a <span class="token operator">&gt;</span> <span class="token number">3</span><span class="token punctuation">:</span>
        <span class="token keyword">return</span> a
    <span class="token keyword">else</span><span class="token punctuation">:</span>
        <span class="token keyword">return</span> <span class="token operator">-</span><span class="token number">1</span>
    
<span class="token comment"># Types annotations remind us that 'a' can be None.</span>
<span class="token comment"># either handle the Optional case (None), or prevent</span>
<span class="token comment"># ever invoking 'func'</span>
<span class="token keyword">def</span> <span class="token function">func</span><span class="token punctuation">(</span>a<span class="token punctuation">:</span> Optional<span class="token punctuation">[</span><span class="token builtin">int</span><span class="token punctuation">]</span><span class="token punctuation">)</span> <span class="token operator">-</span><span class="token operator">&gt;</span> <span class="token builtin">int</span><span class="token punctuation">:</span>
    <span class="token keyword">if</span> a <span class="token keyword">is</span> <span class="token boolean">None</span> <span class="token keyword">or</span> a <span class="token operator">&lt;</span> <span class="token number">3</span><span class="token punctuation">:</span>
        <span class="token keyword">return</span> <span class="token operator">-</span><span class="token number">1</span>
    <span class="token keyword">else</span><span class="token punctuation">:</span>
        <span class="token keyword">return</span> a
</code></pre>
<h2>2. Don't use Python for production services</h2>
<p>If your service has SLA's or has multiple downstream services, don't use Python. To have any chance at service stability, you'll need to fully use MyPy, and have some CI hooks that check for type soundness and prevent non type-annotated code from deploying. Your engineers will need to have the MyPy docs open frequently to maintain the type system, and from experience, you'll get a lot of complaints.</p>
<blockquote>
<p>"I know this code works - our test-suite passes, why am I wasting time getting MyPy to stop complaining about an unbound type?"</p>
</blockquote>
<p>If you have to maintain a type system anyways, it's easier to choose a statically typed language where the compiler builds in type-checking. Also, most popular IDE's for statically typed languages have type hints - so by choosing Python you're just increasing the cognitive load for developers.</p>
<h2>3. Configure MyPy to prevent returning Any</h2>
<p>Developers are people, and people can be lazy. The laziest way to get around MyPy errors is to return type <code>Any</code> - the supertype of all types. This is like returning <code>Any</code> in Scala; as a last resort developers will do this to satisfy the type-checker so they can keep moving. Stop them, or at least make them annotate in code every time they really mean to return <code>Any</code>. Keep the MyPy file as concise as possible, but fail on any errors reported.</p>
<pre class="language-markdown"><code class="language-markdown"><span class="token title important"><span class="token punctuation">#</span> file: mypy.ini</span>

<span class="token title important"><span class="token punctuation">#</span> global options:</span>

[mypy]
warn_return_any = True
</code></pre>
<h1>Takeways</h1>
<p>There's truth to the criticisms of trying to use Python in production. I've personally found that types lead to better code, and try to avoid dynamically typed code without type annotations at all costs - the extra time it takes to reconcile types always leads to stronger production code. It isn't static vs. dynamic, but instead interpreted vs. compiled that is the main differentiator in developer productivity.</p>
<h1>References</h1>
<p>[1] Types for anyone who knows a programming language: <a href="https://www.destroyallsoftware.com/compendium/types?share_key=baf6b67369843fa2">https://www.destroyallsoftware.com/compendium/types?share_key=baf6b67369843fa2</a></p>]]></content:encoded>
            <author>gmgprivacy@proton.me (Matt Goodman)</author>
        </item>
        <item>
            <title><![CDATA[Understanding Polling and Its Failures]]></title>
            <link>https://goodmattg.xyz/articles/Understanding-Polling-and-Its-Failures</link>
            <guid>https://goodmattg.xyz/articles/Understanding-Polling-and-Its-Failures</guid>
            <pubDate>Thu, 28 Sep 2017 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<blockquote>
<p>"Fighting with a large army under your command is nowise different from fighting with a small one: it is merely a question of instituting signs and signals."</p>
<p>-Sun Tzu</p>
</blockquote>
<h2>Overview and Common Assumptions</h2>
<p>The tactical lesson of the 2016 presidential election is that only a fool would blindly trust the polls. Listening to both sides' campaign managers the night before the election it was clear that paradoxically both Trump and Clinton were going to win the election. Somehow the same polls in the same states told the universe of strategists, spectators, wonks, and traders that election result margins were both "stable" and "erroneous", both "decisive" and "within the margin of error". Not to mention the classic campaign strategist spin that "internal polls have us ahead". I contend that media organizations and private companies have spent an inordinate amount of time, money, and effort trying to estimate the result of an election given a collection of polls. These individual polls are known to be flawed, but the reasoning goes that we can use statistics to adjust for all of the errors and estimate the right result. This reasoning is fast becoming an academic relic. Traditional land lines are disappearing are disappearing from American households and response rates are falling across all mediums of communication both traditional and non-traditional [3]. Those organizations that fail to adjust to this new reality will make incorrect predictions, and, more importantly, will provide incorrect analytics. Our group advocates a new estimation methodology and model, discussed in Appendix A, that centers on the assertion that we should estimate election results using larger and therefore more accurate polls from fewer counties.</p>
<p>I am not writing to provide the academic evidence nor the quantitative rational for every one of my statements. I will paint in broad strokes where I find it necessary, dig into the numbers when it is useful, and will offload nuanced in depth explanations to cited works.</p>
<p>Section 3 dives into the mathematical basis of polling. In section 7 I will describe the new polling model our team built and tested over the last two years.</p>
<h2>Disclaimers</h2>
<p>A key assumption of the model we now advocate is that American elections are binary races where voters have a choice between two candidates. Yes, the United States does indeed have more than two political parties. There are the Democrats, Republicans, Green Party, Libertarian, Socialists, etc. When more than 90% of Americans go to the ballot box they vote for either a Democrat or a Republican. Anyone else is an independent candidate - and independents do not win elections in the winner-take-all American system. The highest vote percentage ever achieved by an independent candidate was 19% by Ross Perot in 1992. If an independent candidate doesn't declare early, have significant momentum going into the election, or have a large campaign war chest, there are only two choices for president. Either a Democrat or Republican will win. This is the reason of course that Bernie Sanders joined the Democratic Party. The Independent candidate will play spoiler to the party whom he/she steals more votes from. It is a grim truth that we as Americans are reduced to picking between two parties that are meant to capture every aspect of our beliefs with patchwork platforms appearing to broad coalitions of voters, but it is the truth, and until this system changes we will continue to operate our model assuming two parties.</p>
<p>To my own chagrin, the conservative attack that the Media is biased contains shades of the truth. What the public seems to forget is that the television Media (CNN, MSNBC, FOX) is in the business of selling advertisement slots. While it may have been true that the first television news programs were run as a public service, this is no longer the case. There are very few news programs that offer unbiased news, if such a thing can even exist. CSPAN, ProPublica, and PBS are arguably excellent examples of nonpartisan news. CSPAN is also dry as dirt, boring and substantive like plain oatmeal. Americans want to see action! We like our politics like we like our sports: fast, exciting, and close until the end. No one likes the game that ends 10-0. We want the shoot-out. We want the nail-biter. It's hurts our democracy when the Media chooses to treat politics like sports, but they do it because that's what the market demands.</p>
<h2>Polling Math</h2>
<p>This section should be seen as medicine. I wrote it to be accessible to the average reader, and I hope you take the time to skim it to at least understand that polling isn't all hand-waving and voodoo.</p>
<h3>Expectation of Random Variable</h3>
<p>We have a random variable <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>X</mi></mrow><annotation encoding="application/x-tex">X</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07847em">X</span></span></span></span>. The random variable <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>X</mi></mrow><annotation encoding="application/x-tex">X</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07847em">X</span></span></span></span> has expectation <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="double-struck">E</mi><mo stretchy="false">[</mo><mi>X</mi><mo stretchy="false">]</mo><mo>=</mo><mi>μ</mi></mrow><annotation encoding="application/x-tex">\mathbb{E}[X] = \mu</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mord mathbb">E</span><span class="mopen">[</span><span class="mord mathnormal" style="margin-right:0.07847em">X</span><span class="mclose">]</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.625em;vertical-align:-0.1944em"></span><span class="mord mathnormal">μ</span></span></span></span>. We use random variables to model random phenomena, like rolling dice, or flipping coins. From now on we will use the abbreviation "R.V." interchangeably with random variable. We compute the expectation in the finite case as follows:</p>
<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mi mathvariant="double-struck">E</mi><mo stretchy="false">[</mo><mi>X</mi><mo stretchy="false">]</mo><mo>=</mo><munderover><mo>∑</mo><mrow><mi>i</mi><mo>=</mo><mn>1</mn></mrow><mi mathvariant="normal">∞</mi></munderover><msub><mi>x</mi><mi>i</mi></msub><msub><mi>p</mi><mi>i</mi></msub></mrow><annotation encoding="application/x-tex">\mathbb{E}[X] = \sum_{i=1}^{\infty} x_i p_i</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mord mathbb">E</span><span class="mopen">[</span><span class="mord mathnormal" style="margin-right:0.07847em">X</span><span class="mclose">]</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:2.9291em;vertical-align:-1.2777em"></span><span class="mop op-limits"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.6514em"><span style="top:-1.8723em;margin-left:0em"><span class="pstrut" style="height:3.05em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">i</span><span class="mrel mtight">=</span><span class="mord mtight">1</span></span></span></span><span style="top:-3.05em"><span class="pstrut" style="height:3.05em"></span><span><span class="mop op-symbol large-op">∑</span></span></span><span style="top:-4.3em;margin-left:0em"><span class="pstrut" style="height:3.05em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">∞</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:1.2777em"><span></span></span></span></span></span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord"><span class="mord mathnormal">x</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">i</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em"><span></span></span></span></span></span></span><span class="mord"><span class="mord mathnormal">p</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">i</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em"><span></span></span></span></span></span></span></span></span></span></span>
<p>Don't stop reading, this is actually an easy concept to grasp. The expectation or "expected value" of the random variable <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>X</mi></mrow><annotation encoding="application/x-tex">X</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07847em">X</span></span></span></span> is calculated by adding up every possible value <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>X</mi></mrow><annotation encoding="application/x-tex">X</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07847em">X</span></span></span></span> can take on, multiplied by the probability of <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>X</mi></mrow><annotation encoding="application/x-tex">X</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07847em">X</span></span></span></span> taking on that value. Let's say the random variable <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>X</mi></mrow><annotation encoding="application/x-tex">X</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07847em">X</span></span></span></span> is modeling a dice roll. You and I both know when we roll a six-sided dice it will land one of the values: <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo stretchy="false">{</mo><mn>1</mn><mo separator="true">,</mo><mn>2</mn><mo separator="true">,</mo><mn>3</mn><mo separator="true">,</mo><mn>4</mn><mo separator="true">,</mo><mn>5</mn><mo separator="true">,</mo><mn>6</mn><mo stretchy="false">}</mo></mrow><annotation encoding="application/x-tex">\{1, 2, 3, 4, 5, 6\}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mopen">{</span><span class="mord">1</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord">2</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord">3</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord">4</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord">5</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord">6</span><span class="mclose">}</span></span></span></span>. Assuming no one is cheating, we intuit that there is an equal chance of the dice showing one of the numbers. Using the formula above we an calculate the expectation of the random variable <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>X</mi></mrow><annotation encoding="application/x-tex">X</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07847em">X</span></span></span></span>.</p>
<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mfrac><mn>1</mn><mn>6</mn></mfrac><mo>⋅</mo><mn>1</mn><mo>+</mo><mfrac><mn>1</mn><mn>6</mn></mfrac><mo>⋅</mo><mn>2</mn><mo>+</mo><mfrac><mn>1</mn><mn>6</mn></mfrac><mo>⋅</mo><mn>3</mn><mo>+</mo><mfrac><mn>1</mn><mn>6</mn></mfrac><mo>⋅</mo><mn>4</mn><mo>+</mo><mfrac><mn>1</mn><mn>6</mn></mfrac><mo>⋅</mo><mn>5</mn><mo>+</mo><mfrac><mn>1</mn><mn>6</mn></mfrac><mo>⋅</mo><mn>6</mn><mo>=</mo><mn>3.5</mn></mrow><annotation encoding="application/x-tex">\frac{1}{6} \cdot 1 + \frac{1}{6} \cdot 2 + \frac{1}{6} \cdot 3 + \frac{1}{6} \cdot 4 + \frac{1}{6} \cdot 5 + \frac{1}{6} \cdot 6 = 3.5</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:2.0074em;vertical-align:-0.686em"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.3214em"><span style="top:-2.314em"><span class="pstrut" style="height:3em"></span><span class="mord"><span class="mord">6</span></span></span><span style="top:-3.23em"><span class="pstrut" style="height:3em"></span><span class="frac-line" style="border-bottom-width:0.04em"></span></span><span style="top:-3.677em"><span class="pstrut" style="height:3em"></span><span class="mord"><span class="mord">1</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.686em"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span><span class="mspace" style="margin-right:0.2222em"></span><span class="mbin">⋅</span><span class="mspace" style="margin-right:0.2222em"></span></span><span class="base"><span class="strut" style="height:0.7278em;vertical-align:-0.0833em"></span><span class="mord">1</span><span class="mspace" style="margin-right:0.2222em"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em"></span></span><span class="base"><span class="strut" style="height:2.0074em;vertical-align:-0.686em"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.3214em"><span style="top:-2.314em"><span class="pstrut" style="height:3em"></span><span class="mord"><span class="mord">6</span></span></span><span style="top:-3.23em"><span class="pstrut" style="height:3em"></span><span class="frac-line" style="border-bottom-width:0.04em"></span></span><span style="top:-3.677em"><span class="pstrut" style="height:3em"></span><span class="mord"><span class="mord">1</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.686em"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span><span class="mspace" style="margin-right:0.2222em"></span><span class="mbin">⋅</span><span class="mspace" style="margin-right:0.2222em"></span></span><span class="base"><span class="strut" style="height:0.7278em;vertical-align:-0.0833em"></span><span class="mord">2</span><span class="mspace" style="margin-right:0.2222em"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em"></span></span><span class="base"><span class="strut" style="height:2.0074em;vertical-align:-0.686em"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.3214em"><span style="top:-2.314em"><span class="pstrut" style="height:3em"></span><span class="mord"><span class="mord">6</span></span></span><span style="top:-3.23em"><span class="pstrut" style="height:3em"></span><span class="frac-line" style="border-bottom-width:0.04em"></span></span><span style="top:-3.677em"><span class="pstrut" style="height:3em"></span><span class="mord"><span class="mord">1</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.686em"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span><span class="mspace" style="margin-right:0.2222em"></span><span class="mbin">⋅</span><span class="mspace" style="margin-right:0.2222em"></span></span><span class="base"><span class="strut" style="height:0.7278em;vertical-align:-0.0833em"></span><span class="mord">3</span><span class="mspace" style="margin-right:0.2222em"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em"></span></span><span class="base"><span class="strut" style="height:2.0074em;vertical-align:-0.686em"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.3214em"><span style="top:-2.314em"><span class="pstrut" style="height:3em"></span><span class="mord"><span class="mord">6</span></span></span><span style="top:-3.23em"><span class="pstrut" style="height:3em"></span><span class="frac-line" style="border-bottom-width:0.04em"></span></span><span style="top:-3.677em"><span class="pstrut" style="height:3em"></span><span class="mord"><span class="mord">1</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.686em"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span><span class="mspace" style="margin-right:0.2222em"></span><span class="mbin">⋅</span><span class="mspace" style="margin-right:0.2222em"></span></span><span class="base"><span class="strut" style="height:0.7278em;vertical-align:-0.0833em"></span><span class="mord">4</span><span class="mspace" style="margin-right:0.2222em"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em"></span></span><span class="base"><span class="strut" style="height:2.0074em;vertical-align:-0.686em"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.3214em"><span style="top:-2.314em"><span class="pstrut" style="height:3em"></span><span class="mord"><span class="mord">6</span></span></span><span style="top:-3.23em"><span class="pstrut" style="height:3em"></span><span class="frac-line" style="border-bottom-width:0.04em"></span></span><span style="top:-3.677em"><span class="pstrut" style="height:3em"></span><span class="mord"><span class="mord">1</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.686em"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span><span class="mspace" style="margin-right:0.2222em"></span><span class="mbin">⋅</span><span class="mspace" style="margin-right:0.2222em"></span></span><span class="base"><span class="strut" style="height:0.7278em;vertical-align:-0.0833em"></span><span class="mord">5</span><span class="mspace" style="margin-right:0.2222em"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em"></span></span><span class="base"><span class="strut" style="height:2.0074em;vertical-align:-0.686em"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.3214em"><span style="top:-2.314em"><span class="pstrut" style="height:3em"></span><span class="mord"><span class="mord">6</span></span></span><span style="top:-3.23em"><span class="pstrut" style="height:3em"></span><span class="frac-line" style="border-bottom-width:0.04em"></span></span><span style="top:-3.677em"><span class="pstrut" style="height:3em"></span><span class="mord"><span class="mord">1</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.686em"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span><span class="mspace" style="margin-right:0.2222em"></span><span class="mbin">⋅</span><span class="mspace" style="margin-right:0.2222em"></span></span><span class="base"><span class="strut" style="height:0.6444em"></span><span class="mord">6</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.6444em"></span><span class="mord">3.5</span></span></span></span></span>
<p>I might not get around to showing the proof, but something called The Law of Large Numbers says that if we look at the value the random variable spits out many, many times, the difference between the average of those spit out values and the expectation we computed above will fall to <strong>zero</strong>. This needs to be very clear. I roll the dice once and it turns up a 3. The second time I roll it turns up a 2. The next time a 5. The next time 5. If I roll the dice a billion times, add up all the numbers I roll, and divide by a billion, the answer is all but guaranteed to be incredibly close to 3.5.</p>
<h3>Markov's Inequality</h3>
<p>So what was the point of this? We're getting to the good part. Here's the question you should be asking (that I borrowed [1]):</p>
<blockquote>
<p>"What is the probability that the value of the random variable X, is not close to its expectation?"</p>
</blockquote>
<p>I just told you that we can take a random variable and compute the value we expect as the average if we can 'roll the dice' an infinite number of times. But in the real world, you don't get to roll the dice billions of times, you often just get one roll of the dice. So we want to know how likely it is that the random variable spits out a value that isn't close at all to the expectation. First I'll give the formula, then we'll unpack it together, finally I'll give the proof.</p>
<p>Markov's Inequality: For a nonnegative random variable, <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>X</mi><mo>:</mo><mi mathvariant="normal">Ω</mi><mo>→</mo><mi mathvariant="double-struck">R</mi></mrow><annotation encoding="application/x-tex">X: \Omega \rightarrow \mathbb{R}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07847em">X</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">:</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord">Ω</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">→</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.6889em"></span><span class="mord mathbb">R</span></span></span></span>, where <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>X</mi><mo stretchy="false">(</mo><mi>s</mi><mo stretchy="false">)</mo><mo>≥</mo><mn>0</mn></mrow><annotation encoding="application/x-tex">X(s) \geq 0</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mord mathnormal" style="margin-right:0.07847em">X</span><span class="mopen">(</span><span class="mord mathnormal">s</span><span class="mclose">)</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">≥</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.6444em"></span><span class="mord">0</span></span></span></span> for all <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>s</mi><mo>∈</mo><mi mathvariant="normal">Ω</mi></mrow><annotation encoding="application/x-tex">s \in \Omega</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.5782em;vertical-align:-0.0391em"></span><span class="mord mathnormal">s</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">∈</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord">Ω</span></span></span></span>, for any positive real number <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>a</mi><mo>&gt;</mo><mn>0</mn></mrow><annotation encoding="application/x-tex">a &gt; 0</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.5782em;vertical-align:-0.0391em"></span><span class="mord mathnormal">a</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">&gt;</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.6444em"></span><span class="mord">0</span></span></span></span>:</p>
<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mi>P</mi><mo stretchy="false">(</mo><mi>X</mi><mo>≥</mo><mi>a</mi><mo stretchy="false">)</mo><mo>≤</mo><mfrac><mrow><mi mathvariant="double-struck">E</mi><mo stretchy="false">[</mo><mi>X</mi><mo stretchy="false">]</mo></mrow><mi>a</mi></mfrac></mrow><annotation encoding="application/x-tex">P(X \geq a) \leq \frac{\mathbb{E}[X]}{a}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mord mathnormal" style="margin-right:0.13889em">P</span><span class="mopen">(</span><span class="mord mathnormal" style="margin-right:0.07847em">X</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">≥</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mord mathnormal">a</span><span class="mclose">)</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">≤</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:2.113em;vertical-align:-0.686em"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.427em"><span style="top:-2.314em"><span class="pstrut" style="height:3em"></span><span class="mord"><span class="mord mathnormal">a</span></span></span><span style="top:-3.23em"><span class="pstrut" style="height:3em"></span><span class="frac-line" style="border-bottom-width:0.04em"></span></span><span style="top:-3.677em"><span class="pstrut" style="height:3em"></span><span class="mord"><span class="mord mathbb">E</span><span class="mopen">[</span><span class="mord mathnormal" style="margin-right:0.07847em">X</span><span class="mclose">]</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.686em"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span></span></span></span></span>
<p><strong>Unpacking the formula</strong></p>
<p>As promised, let's unpack the formula step-by-step. We already know that <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>X</mi></mrow><annotation encoding="application/x-tex">X</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07847em">X</span></span></span></span> is a random variable, but what does <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="normal">Ω</mi><mo>→</mo><mi mathvariant="double-struck">R</mi></mrow><annotation encoding="application/x-tex">\Omega \rightarrow \mathbb{R}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord">Ω</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">→</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.6889em"></span><span class="mord mathbb">R</span></span></span></span> mean? <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="normal">Ω</mi></mrow><annotation encoding="application/x-tex">\Omega</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord">Ω</span></span></span></span> is a symbol that means "the probability space". So whenever you see the symbol <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="normal">Ω</mi></mrow><annotation encoding="application/x-tex">\Omega</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord">Ω</span></span></span></span>, substitute the words "the probability space" in. Every time you see <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>P</mi></mrow><annotation encoding="application/x-tex">P</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.13889em">P</span></span></span></span> and then parentheses, substitute the words "the probability of". <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>P</mi><mo stretchy="false">(</mo><mtext>dessert&nbsp;for&nbsp;dinner</mtext><mo stretchy="false">)</mo></mrow><annotation encoding="application/x-tex">P(\text{dessert for dinner})</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mord mathnormal" style="margin-right:0.13889em">P</span><span class="mopen">(</span><span class="mord text"><span class="mord">dessert&nbsp;for&nbsp;dinner</span></span><span class="mclose">)</span></span></span></span> is the same as "the probability I will eat dessert for dinner".</p>
<p>The probability space is the space of everything that can happen in our probability world. In the world where <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>X</mi></mrow><annotation encoding="application/x-tex">X</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07847em">X</span></span></span></span> represents a dice roll, <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="normal">Ω</mi></mrow><annotation encoding="application/x-tex">\Omega</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord">Ω</span></span></span></span> means every integer between <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo>−</mo><mi mathvariant="normal">∞</mi></mrow><annotation encoding="application/x-tex">-\infty</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em"></span><span class="mord">−</span><span class="mord">∞</span></span></span></span> to <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="normal">∞</mi></mrow><annotation encoding="application/x-tex">\infty</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em"></span><span class="mord">∞</span></span></span></span>. This means that <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="normal">Ω</mi></mrow><annotation encoding="application/x-tex">\Omega</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord">Ω</span></span></span></span> includes the numbers 1, -57, 263, -1454, and so on. "But if I roll a dice, the number -57 will never show up?!". That's right! The symbols <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="normal">Ω</mi><mo>→</mo><mi mathvariant="double-struck">R</mi></mrow><annotation encoding="application/x-tex">\Omega \rightarrow \mathbb{R}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord">Ω</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">→</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.6889em"></span><span class="mord mathbb">R</span></span></span></span> means that <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>X</mi></mrow><annotation encoding="application/x-tex">X</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07847em">X</span></span></span></span> is actually a function. It takes in some event from the probability space <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="normal">Ω</mi></mrow><annotation encoding="application/x-tex">\Omega</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord">Ω</span></span></span></span> and it spits out a real number <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="double-struck">R</mi></mrow><annotation encoding="application/x-tex">\mathbb{R}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6889em"></span><span class="mord mathbb">R</span></span></span></span> that is greater than or equal to zero. So in the the case where <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>X</mi></mrow><annotation encoding="application/x-tex">X</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07847em">X</span></span></span></span> is a dice roll, here's a few examples of what <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>X</mi><mo>:</mo><mi mathvariant="normal">Ω</mi><mo>→</mo><mi mathvariant="double-struck">R</mi></mrow><annotation encoding="application/x-tex">X: \Omega \rightarrow \mathbb{R}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07847em">X</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">:</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord">Ω</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">→</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.6889em"></span><span class="mord mathbb">R</span></span></span></span> spits out.</p>
<p>The probability that <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>X</mi></mrow><annotation encoding="application/x-tex">X</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07847em">X</span></span></span></span> takes on the value '1' is <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mfrac><mn>1</mn><mn>6</mn></mfrac></mrow><annotation encoding="application/x-tex">\frac{1}{6}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1.1901em;vertical-align:-0.345em"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.8451em"><span style="top:-2.655em"><span class="pstrut" style="height:3em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">6</span></span></span></span><span style="top:-3.23em"><span class="pstrut" style="height:3em"></span><span class="frac-line" style="border-bottom-width:0.04em"></span></span><span style="top:-3.394em"><span class="pstrut" style="height:3em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">1</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.345em"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span></span></span></span>:</p>
<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mi>P</mi><mo stretchy="false">(</mo><mi>X</mi><mo>=</mo><mn>1</mn><mo stretchy="false">)</mo><mo>=</mo><mfrac><mn>1</mn><mn>6</mn></mfrac></mrow><annotation encoding="application/x-tex">P(X = 1) = \frac{1}{6}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mord mathnormal" style="margin-right:0.13889em">P</span><span class="mopen">(</span><span class="mord mathnormal" style="margin-right:0.07847em">X</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mord">1</span><span class="mclose">)</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:2.0074em;vertical-align:-0.686em"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.3214em"><span style="top:-2.314em"><span class="pstrut" style="height:3em"></span><span class="mord"><span class="mord">6</span></span></span><span style="top:-3.23em"><span class="pstrut" style="height:3em"></span><span class="frac-line" style="border-bottom-width:0.04em"></span></span><span style="top:-3.677em"><span class="pstrut" style="height:3em"></span><span class="mord"><span class="mord">1</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.686em"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span></span></span></span></span>
<p>The probability that <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>X</mi></mrow><annotation encoding="application/x-tex">X</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07847em">X</span></span></span></span> takes on the value '-57' is zero. It will never happen:</p>
<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mi>P</mi><mo stretchy="false">(</mo><mi>X</mi><mo>=</mo><mo>−</mo><mn>57</mn><mo stretchy="false">)</mo><mo>=</mo><mn>0</mn></mrow><annotation encoding="application/x-tex">P(X = -57) = 0</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mord mathnormal" style="margin-right:0.13889em">P</span><span class="mopen">(</span><span class="mord mathnormal" style="margin-right:0.07847em">X</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mord">−</span><span class="mord">57</span><span class="mclose">)</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.6444em"></span><span class="mord">0</span></span></span></span></span>
<p>The probability that <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>X</mi></mrow><annotation encoding="application/x-tex">X</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07847em">X</span></span></span></span> takes on the value '4' is <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mfrac><mn>1</mn><mn>6</mn></mfrac></mrow><annotation encoding="application/x-tex">\frac{1}{6}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1.1901em;vertical-align:-0.345em"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.8451em"><span style="top:-2.655em"><span class="pstrut" style="height:3em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">6</span></span></span></span><span style="top:-3.23em"><span class="pstrut" style="height:3em"></span><span class="frac-line" style="border-bottom-width:0.04em"></span></span><span style="top:-3.394em"><span class="pstrut" style="height:3em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">1</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.345em"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span></span></span></span>:</p>
<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mi>P</mi><mo stretchy="false">(</mo><mi>X</mi><mo>=</mo><mn>4</mn><mo stretchy="false">)</mo><mo>=</mo><mfrac><mn>1</mn><mn>6</mn></mfrac></mrow><annotation encoding="application/x-tex">P(X = 4) = \frac{1}{6}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mord mathnormal" style="margin-right:0.13889em">P</span><span class="mopen">(</span><span class="mord mathnormal" style="margin-right:0.07847em">X</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mord">4</span><span class="mclose">)</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:2.0074em;vertical-align:-0.686em"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.3214em"><span style="top:-2.314em"><span class="pstrut" style="height:3em"></span><span class="mord"><span class="mord">6</span></span></span><span style="top:-3.23em"><span class="pstrut" style="height:3em"></span><span class="frac-line" style="border-bottom-width:0.04em"></span></span><span style="top:-3.677em"><span class="pstrut" style="height:3em"></span><span class="mord"><span class="mord">1</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.686em"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span></span></span></span></span>
<p>The probability that <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>X</mi></mrow><annotation encoding="application/x-tex">X</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07847em">X</span></span></span></span> takes on the value '7' is zero. It will never happen:</p>
<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mi>P</mi><mo stretchy="false">(</mo><mi>X</mi><mo>=</mo><mn>7</mn><mo stretchy="false">)</mo><mo>=</mo><mn>0</mn></mrow><annotation encoding="application/x-tex">P(X = 7) = 0</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mord mathnormal" style="margin-right:0.13889em">P</span><span class="mopen">(</span><span class="mord mathnormal" style="margin-right:0.07847em">X</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mord">7</span><span class="mclose">)</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.6444em"></span><span class="mord">0</span></span></span></span></span>
<p>So that's pretty reasonable. <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>X</mi></mrow><annotation encoding="application/x-tex">X</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07847em">X</span></span></span></span> takes in an event that belongs to the complete probability space <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="normal">Ω</mi></mrow><annotation encoding="application/x-tex">\Omega</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord">Ω</span></span></span></span> and spits out a number that is either zero, meaning the event will never happen, one, meaning the event will literally always occur, or something between zero and one, meaning the event has some probability of occurring. Got it? Onwards then!</p>
<p><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>a</mi></mrow><annotation encoding="application/x-tex">a</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em"></span><span class="mord mathnormal">a</span></span></span></span> is a constant that is greater than 0. We get to decide what the constant <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>a</mi></mrow><annotation encoding="application/x-tex">a</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em"></span><span class="mord mathnormal">a</span></span></span></span> is. You want <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>a</mi></mrow><annotation encoding="application/x-tex">a</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em"></span><span class="mord mathnormal">a</span></span></span></span> to be 0.00001. Okay. You want <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>a</mi></mrow><annotation encoding="application/x-tex">a</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em"></span><span class="mord mathnormal">a</span></span></span></span> to be 100000. Okay. As long as <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>a</mi></mrow><annotation encoding="application/x-tex">a</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em"></span><span class="mord mathnormal">a</span></span></span></span> isn't less than or equal to zero, you're all good. We'll see later that we can pick specific values of <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>a</mi></mrow><annotation encoding="application/x-tex">a</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em"></span><span class="mord mathnormal">a</span></span></span></span> that let us do interesting things, but for now just remember that we decide what value <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>a</mi></mrow><annotation encoding="application/x-tex">a</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em"></span><span class="mord mathnormal">a</span></span></span></span> takes on.</p>
<p>That's all there is to it! We know what everything means. The following section is the derivation of Chebyshev's Inequality from Markov's Inequality.</p>
<p>First the proof of Markov's Inequality (taken from [1]):</p>
<p>Proof. Let the event <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>A</mi><mo>⊆</mo><mi mathvariant="normal">Ω</mi></mrow><annotation encoding="application/x-tex">A \subseteq \Omega</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8193em;vertical-align:-0.136em"></span><span class="mord mathnormal">A</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">⊆</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord">Ω</span></span></span></span> be defined by: <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>A</mi><mo>=</mo><mo stretchy="false">{</mo><mi>s</mi><mo>∈</mo><mi mathvariant="normal">Ω</mi><mo>∣</mo><mi>X</mi><mo stretchy="false">(</mo><mi>s</mi><mo stretchy="false">)</mo><mo>≥</mo><mi>a</mi><mo stretchy="false">}</mo></mrow><annotation encoding="application/x-tex">A = \{s \in \Omega \mid X(s) \geq a\}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal">A</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mopen">{</span><span class="mord mathnormal">s</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">∈</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mord">Ω</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">∣</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mord mathnormal" style="margin-right:0.07847em">X</span><span class="mopen">(</span><span class="mord mathnormal">s</span><span class="mclose">)</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">≥</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mord mathnormal">a</span><span class="mclose">}</span></span></span></span>. We want to prove <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>P</mi><mo stretchy="false">(</mo><mi>A</mi><mo stretchy="false">)</mo><mo>≤</mo><mfrac><mrow><mi mathvariant="double-struck">E</mi><mo stretchy="false">[</mo><mi>X</mi><mo stretchy="false">]</mo></mrow><mi>a</mi></mfrac></mrow><annotation encoding="application/x-tex">P(A) \leq \frac{\mathbb{E}[X]}{a}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mord mathnormal" style="margin-right:0.13889em">P</span><span class="mopen">(</span><span class="mord mathnormal">A</span><span class="mclose">)</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">≤</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:1.355em;vertical-align:-0.345em"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.01em"><span style="top:-2.655em"><span class="pstrut" style="height:3em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">a</span></span></span></span><span style="top:-3.23em"><span class="pstrut" style="height:3em"></span><span class="frac-line" style="border-bottom-width:0.04em"></span></span><span style="top:-3.485em"><span class="pstrut" style="height:3em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathbb mtight">E</span><span class="mopen mtight">[</span><span class="mord mathnormal mtight" style="margin-right:0.07847em">X</span><span class="mclose mtight">]</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.345em"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span></span></span></span>.</p>
<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mtable rowspacing="0.25em" columnalign="right left right left" columnspacing="0em 1em 0em"><mtr><mtd><mstyle scriptlevel="0" displaystyle="true"><mrow><mi mathvariant="double-struck">E</mi><mo stretchy="false">[</mo><mi>X</mi><mo stretchy="false">]</mo></mrow></mstyle></mtd><mtd><mstyle scriptlevel="0" displaystyle="true"><mrow><mrow></mrow><mo>=</mo><munder><mo>∑</mo><mrow><mi>s</mi><mo>∈</mo><mi mathvariant="normal">Ω</mi></mrow></munder><mi>P</mi><mo stretchy="false">(</mo><mi>s</mi><mo stretchy="false">)</mo><mi>X</mi><mo stretchy="false">(</mo><mi>s</mi><mo stretchy="false">)</mo></mrow></mstyle></mtd></mtr><mtr><mtd><mstyle scriptlevel="0" displaystyle="true"><mrow></mrow></mstyle></mtd><mtd><mstyle scriptlevel="0" displaystyle="true"><mrow><mrow></mrow><mo>=</mo><munder><mo>∑</mo><mrow><mi>s</mi><mo>∈</mo><mi>A</mi></mrow></munder><mi>P</mi><mo stretchy="false">(</mo><mi>s</mi><mo stretchy="false">)</mo><mi>X</mi><mo stretchy="false">(</mo><mi>s</mi><mo stretchy="false">)</mo><mo>+</mo><munder><mo>∑</mo><mrow><mi>s</mi><mo mathvariant="normal">∉</mo><mi>A</mi></mrow></munder><mi>P</mi><mo stretchy="false">(</mo><mi>s</mi><mo stretchy="false">)</mo><mi>X</mi><mo stretchy="false">(</mo><mi>s</mi><mo stretchy="false">)</mo></mrow></mstyle></mtd></mtr><mtr><mtd><mstyle scriptlevel="0" displaystyle="true"><mrow></mrow></mstyle></mtd><mtd><mstyle scriptlevel="0" displaystyle="true"><mrow><mrow></mrow><mo>≥</mo><munder><mo>∑</mo><mrow><mi>s</mi><mo>∈</mo><mi>A</mi></mrow></munder><mi>P</mi><mo stretchy="false">(</mo><mi>s</mi><mo stretchy="false">)</mo><mi>X</mi><mo stretchy="false">(</mo><mi>s</mi><mo stretchy="false">)</mo></mrow></mstyle></mtd><mtd><mstyle scriptlevel="0" displaystyle="true"><mrow></mrow></mstyle></mtd><mtd><mstyle scriptlevel="0" displaystyle="true"><mrow><mrow></mrow><mtext>(because&nbsp;</mtext><mi>X</mi><mo stretchy="false">(</mo><mi>s</mi><mo stretchy="false">)</mo><mo>≥</mo><mi>a</mi><mtext>&nbsp;for&nbsp;all&nbsp;</mtext><mi>s</mi><mo>∈</mo><mi mathvariant="normal">Ω</mi><mtext>)</mtext></mrow></mstyle></mtd></mtr><mtr><mtd><mstyle scriptlevel="0" displaystyle="true"><mrow></mrow></mstyle></mtd><mtd><mstyle scriptlevel="0" displaystyle="true"><mrow><mrow></mrow><mo>≥</mo><munder><mo>∑</mo><mrow><mi>s</mi><mo>∈</mo><mi>A</mi></mrow></munder><mi>P</mi><mo stretchy="false">(</mo><mi>s</mi><mo stretchy="false">)</mo><mi>a</mi></mrow></mstyle></mtd><mtd><mstyle scriptlevel="0" displaystyle="true"><mrow></mrow></mstyle></mtd><mtd><mstyle scriptlevel="0" displaystyle="true"><mrow><mrow></mrow><mtext>(because&nbsp;</mtext><mi>X</mi><mo stretchy="false">(</mo><mi>s</mi><mo stretchy="false">)</mo><mo>≥</mo><mi>a</mi><mtext>&nbsp;for&nbsp;all&nbsp;</mtext><mi>s</mi><mo>∈</mo><mi>A</mi><mtext>)</mtext></mrow></mstyle></mtd></mtr><mtr><mtd><mstyle scriptlevel="0" displaystyle="true"><mrow></mrow></mstyle></mtd><mtd><mstyle scriptlevel="0" displaystyle="true"><mrow><mrow></mrow><mo>=</mo><mi>a</mi><munder><mo>∑</mo><mrow><mi>s</mi><mo>∈</mo><mi>A</mi></mrow></munder><mi>P</mi><mo stretchy="false">(</mo><mi>s</mi><mo stretchy="false">)</mo><mo>=</mo><mi>a</mi><mo>⋅</mo><mi>P</mi><mo stretchy="false">(</mo><mi>A</mi><mo stretchy="false">)</mo></mrow></mstyle></mtd></mtr><mtr><mtd><mstyle scriptlevel="0" displaystyle="true"><mrow><mi mathvariant="double-struck">E</mi><mo stretchy="false">[</mo><mi>X</mi><mo stretchy="false">]</mo></mrow></mstyle></mtd><mtd><mstyle scriptlevel="0" displaystyle="true"><mrow><mrow></mrow><mo>≥</mo><mi>a</mi><mo>⋅</mo><mi>P</mi><mo stretchy="false">(</mo><mi>A</mi><mo stretchy="false">)</mo></mrow></mstyle></mtd></mtr></mtable><annotation encoding="application/x-tex">\begin{aligned}
\mathbb{E}[X] &amp;= \sum_{s \in \Omega} P(s)X(s) \\
&amp;= \sum_{s \in A} P(s)X(s) + \sum_{s \notin A} P(s)X(s) \\
&amp;\geq \sum_{s \in A} P(s)X(s) &amp;&amp; \text{(because } X(s) \geq a \text{ for all } s \in \Omega \text{)} \\
&amp;\geq \sum_{s \in A} P(s)a &amp;&amp; \text{(because } X(s) \geq a \text{ for all } s \in A \text{)} \\
&amp;= a \sum_{s \in A} P(s) = a \cdot P(A) \\
\mathbb{E}[X] &amp;\geq a \cdot P(A)
\end{aligned}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:15.0529em;vertical-align:-7.2764em"></span><span class="mord"><span class="mtable"><span class="col-align-r"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:7.7764em"><span style="top:-9.7764em"><span class="pstrut" style="height:3.05em"></span><span class="mord"><span class="mord mathbb">E</span><span class="mopen">[</span><span class="mord mathnormal" style="margin-right:0.07847em">X</span><span class="mclose">]</span></span></span><span style="top:-7.1047em"><span class="pstrut" style="height:3.05em"></span><span class="mord"></span></span><span style="top:-4.2387em"><span class="pstrut" style="height:3.05em"></span><span class="mord"></span></span><span style="top:-1.567em"><span class="pstrut" style="height:3.05em"></span><span class="mord"></span></span><span style="top:1.1047em"><span class="pstrut" style="height:3.05em"></span><span class="mord"></span></span><span style="top:3.5664em"><span class="pstrut" style="height:3.05em"></span><span class="mord"><span class="mord mathbb">E</span><span class="mopen">[</span><span class="mord mathnormal" style="margin-right:0.07847em">X</span><span class="mclose">]</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:7.2764em"><span></span></span></span></span></span><span class="col-align-l"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:7.7764em"><span style="top:-9.7764em"><span class="pstrut" style="height:3.05em"></span><span class="mord"><span class="mord"></span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mop op-limits"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.05em"><span style="top:-1.8557em;margin-left:0em"><span class="pstrut" style="height:3.05em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">s</span><span class="mrel mtight">∈</span><span class="mord mtight">Ω</span></span></span></span><span style="top:-3.05em"><span class="pstrut" style="height:3.05em"></span><span><span class="mop op-symbol large-op">∑</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:1.3217em"><span></span></span></span></span></span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord mathnormal" style="margin-right:0.13889em">P</span><span class="mopen">(</span><span class="mord mathnormal">s</span><span class="mclose">)</span><span class="mord mathnormal" style="margin-right:0.07847em">X</span><span class="mopen">(</span><span class="mord mathnormal">s</span><span class="mclose">)</span></span></span><span style="top:-7.1047em"><span class="pstrut" style="height:3.05em"></span><span class="mord"><span class="mord"></span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mop op-limits"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.05em"><span style="top:-1.8557em;margin-left:0em"><span class="pstrut" style="height:3.05em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">s</span><span class="mrel mtight">∈</span><span class="mord mathnormal mtight">A</span></span></span></span><span style="top:-3.05em"><span class="pstrut" style="height:3.05em"></span><span><span class="mop op-symbol large-op">∑</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:1.3217em"><span></span></span></span></span></span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord mathnormal" style="margin-right:0.13889em">P</span><span class="mopen">(</span><span class="mord mathnormal">s</span><span class="mclose">)</span><span class="mord mathnormal" style="margin-right:0.07847em">X</span><span class="mopen">(</span><span class="mord mathnormal">s</span><span class="mclose">)</span><span class="mspace" style="margin-right:0.2222em"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em"></span><span class="mop op-limits"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.05em"><span style="top:-1.809em;margin-left:0em"><span class="pstrut" style="height:3.05em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">s</span><span class="mrel mtight"><span class="mord mtight"><span class="mrel mtight">∈</span></span><span class="mord vbox mtight"><span class="thinbox mtight"><span class="llap mtight"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="inner"><span class="mord mtight"><span class="mord mtight">/</span><span class="mspace mtight" style="margin-right:0.0651em"></span></span></span><span class="fix"></span></span></span></span></span><span class="mord mathnormal mtight">A</span></span></span></span><span style="top:-3.05em"><span class="pstrut" style="height:3.05em"></span><span><span class="mop op-symbol large-op">∑</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:1.516em"><span></span></span></span></span></span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord mathnormal" style="margin-right:0.13889em">P</span><span class="mopen">(</span><span class="mord mathnormal">s</span><span class="mclose">)</span><span class="mord mathnormal" style="margin-right:0.07847em">X</span><span class="mopen">(</span><span class="mord mathnormal">s</span><span class="mclose">)</span></span></span><span style="top:-4.2387em"><span class="pstrut" style="height:3.05em"></span><span class="mord"><span class="mord"></span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">≥</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mop op-limits"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.05em"><span style="top:-1.8557em;margin-left:0em"><span class="pstrut" style="height:3.05em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">s</span><span class="mrel mtight">∈</span><span class="mord mathnormal mtight">A</span></span></span></span><span style="top:-3.05em"><span class="pstrut" style="height:3.05em"></span><span><span class="mop op-symbol large-op">∑</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:1.3217em"><span></span></span></span></span></span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord mathnormal" style="margin-right:0.13889em">P</span><span class="mopen">(</span><span class="mord mathnormal">s</span><span class="mclose">)</span><span class="mord mathnormal" style="margin-right:0.07847em">X</span><span class="mopen">(</span><span class="mord mathnormal">s</span><span class="mclose">)</span></span></span><span style="top:-1.567em"><span class="pstrut" style="height:3.05em"></span><span class="mord"><span class="mord"></span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">≥</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mop op-limits"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.05em"><span style="top:-1.8557em;margin-left:0em"><span class="pstrut" style="height:3.05em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">s</span><span class="mrel mtight">∈</span><span class="mord mathnormal mtight">A</span></span></span></span><span style="top:-3.05em"><span class="pstrut" style="height:3.05em"></span><span><span class="mop op-symbol large-op">∑</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:1.3217em"><span></span></span></span></span></span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord mathnormal" style="margin-right:0.13889em">P</span><span class="mopen">(</span><span class="mord mathnormal">s</span><span class="mclose">)</span><span class="mord mathnormal">a</span></span></span><span style="top:1.1047em"><span class="pstrut" style="height:3.05em"></span><span class="mord"><span class="mord"></span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mord mathnormal">a</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mop op-limits"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.05em"><span style="top:-1.8557em;margin-left:0em"><span class="pstrut" style="height:3.05em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">s</span><span class="mrel mtight">∈</span><span class="mord mathnormal mtight">A</span></span></span></span><span style="top:-3.05em"><span class="pstrut" style="height:3.05em"></span><span><span class="mop op-symbol large-op">∑</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:1.3217em"><span></span></span></span></span></span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord mathnormal" style="margin-right:0.13889em">P</span><span class="mopen">(</span><span class="mord mathnormal">s</span><span class="mclose">)</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mord mathnormal">a</span><span class="mspace" style="margin-right:0.2222em"></span><span class="mbin">⋅</span><span class="mspace" style="margin-right:0.2222em"></span><span class="mord mathnormal" style="margin-right:0.13889em">P</span><span class="mopen">(</span><span class="mord mathnormal">A</span><span class="mclose">)</span></span></span><span style="top:3.5664em"><span class="pstrut" style="height:3.05em"></span><span class="mord"><span class="mord"></span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">≥</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mord mathnormal">a</span><span class="mspace" style="margin-right:0.2222em"></span><span class="mbin">⋅</span><span class="mspace" style="margin-right:0.2222em"></span><span class="mord mathnormal" style="margin-right:0.13889em">P</span><span class="mopen">(</span><span class="mord mathnormal">A</span><span class="mclose">)</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:7.2764em"><span></span></span></span></span></span><span class="arraycolsep" style="width:1em"></span><span class="col-align-r"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:2.2387em"><span style="top:-4.2387em"><span class="pstrut" style="height:3.05em"></span><span class="mord"></span></span><span style="top:-1.567em"><span class="pstrut" style="height:3.05em"></span><span class="mord"></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:3.1047em"><span></span></span></span></span></span><span class="col-align-l"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:2.2387em"><span style="top:-4.2387em"><span class="pstrut" style="height:3.05em"></span><span class="mord"><span class="mord"></span><span class="mord text"><span class="mord">(because&nbsp;</span></span><span class="mord mathnormal" style="margin-right:0.07847em">X</span><span class="mopen">(</span><span class="mord mathnormal">s</span><span class="mclose">)</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">≥</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mord mathnormal">a</span><span class="mord text"><span class="mord">&nbsp;for&nbsp;all&nbsp;</span></span><span class="mord mathnormal">s</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">∈</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mord">Ω</span><span class="mord text"><span class="mord">)</span></span></span></span><span style="top:-1.567em"><span class="pstrut" style="height:3.05em"></span><span class="mord"><span class="mord"></span><span class="mord text"><span class="mord">(because&nbsp;</span></span><span class="mord mathnormal" style="margin-right:0.07847em">X</span><span class="mopen">(</span><span class="mord mathnormal">s</span><span class="mclose">)</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">≥</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mord mathnormal">a</span><span class="mord text"><span class="mord">&nbsp;for&nbsp;all&nbsp;</span></span><span class="mord mathnormal">s</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">∈</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mord mathnormal">A</span><span class="mord text"><span class="mord">)</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:3.1047em"><span></span></span></span></span></span></span></span></span></span></span></span>
<p>Thus, <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>P</mi><mo stretchy="false">(</mo><mi>X</mi><mo>≥</mo><mi>a</mi><mo stretchy="false">)</mo><mo>≤</mo><mfrac><mrow><mi mathvariant="double-struck">E</mi><mo stretchy="false">[</mo><mi>X</mi><mo stretchy="false">]</mo></mrow><mi>a</mi></mfrac></mrow><annotation encoding="application/x-tex">P(X \geq a) \leq \frac{\mathbb{E}[X]}{a}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mord mathnormal" style="margin-right:0.13889em">P</span><span class="mopen">(</span><span class="mord mathnormal" style="margin-right:0.07847em">X</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">≥</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mord mathnormal">a</span><span class="mclose">)</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">≤</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:1.355em;vertical-align:-0.345em"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.01em"><span style="top:-2.655em"><span class="pstrut" style="height:3em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">a</span></span></span></span><span style="top:-3.23em"><span class="pstrut" style="height:3em"></span><span class="frac-line" style="border-bottom-width:0.04em"></span></span><span style="top:-3.485em"><span class="pstrut" style="height:3em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathbb mtight">E</span><span class="mopen mtight">[</span><span class="mord mathnormal mtight" style="margin-right:0.07847em">X</span><span class="mclose mtight">]</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.345em"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span></span></span></span>.</p>
<p>Markov's Inequality is a general statement that holds for any non-negative random variable. It is useful for defining a rough upper bound on the probability of an event. For the case of polling, remember that we wanted to answer the question of how likely it is that our estimation is far away from the truth. For that we need to use Chebyshev's inequality, which can be defined as a special case of Markov's Inequality.</p>
<p>Chebyshev's Inequality: For any random variable, <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>X</mi><mo>:</mo><mi mathvariant="normal">Ω</mi><mo>→</mo><mi mathvariant="double-struck">R</mi></mrow><annotation encoding="application/x-tex">X: \Omega \rightarrow \mathbb{R}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07847em">X</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">:</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord">Ω</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">→</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.6889em"></span><span class="mord mathbb">R</span></span></span></span>, and let <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>r</mi><mo>&gt;</mo><mn>0</mn></mrow><annotation encoding="application/x-tex">r &gt; 0</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.5782em;vertical-align:-0.0391em"></span><span class="mord mathnormal" style="margin-right:0.02778em">r</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">&gt;</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.6444em"></span><span class="mord">0</span></span></span></span> be any positive real number.</p>
<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mi>P</mi><mo stretchy="false">(</mo><mi mathvariant="normal">∣</mi><mi>X</mi><mo>−</mo><mi mathvariant="double-struck">E</mi><mo stretchy="false">[</mo><mi>X</mi><mo stretchy="false">]</mo><mi mathvariant="normal">∣</mi><mo>≥</mo><mi>r</mi><mo stretchy="false">)</mo><mo>≤</mo><mfrac><mrow><mi>V</mi><mo stretchy="false">(</mo><mi>X</mi><mo stretchy="false">)</mo></mrow><msup><mi>r</mi><mn>2</mn></msup></mfrac></mrow><annotation encoding="application/x-tex">P(|X - \mathbb{E}[X]| \geq r) \leq \frac{V(X)}{r^2}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mord mathnormal" style="margin-right:0.13889em">P</span><span class="mopen">(</span><span class="mord">∣</span><span class="mord mathnormal" style="margin-right:0.07847em">X</span><span class="mspace" style="margin-right:0.2222em"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mord mathbb">E</span><span class="mopen">[</span><span class="mord mathnormal" style="margin-right:0.07847em">X</span><span class="mclose">]</span><span class="mord">∣</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">≥</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mord mathnormal" style="margin-right:0.02778em">r</span><span class="mclose">)</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">≤</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:2.113em;vertical-align:-0.686em"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.427em"><span style="top:-2.314em"><span class="pstrut" style="height:3em"></span><span class="mord"><span class="mord"><span class="mord mathnormal" style="margin-right:0.02778em">r</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.7401em"><span style="top:-2.989em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight">2</span></span></span></span></span></span></span></span></span></span><span style="top:-3.23em"><span class="pstrut" style="height:3em"></span><span class="frac-line" style="border-bottom-width:0.04em"></span></span><span style="top:-3.677em"><span class="pstrut" style="height:3em"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.22222em">V</span><span class="mopen">(</span><span class="mord mathnormal" style="margin-right:0.07847em">X</span><span class="mclose">)</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.686em"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span></span></span></span></span>
<p>Again, let's spend some time unpacking this formula. With Markov's Inequality we could compute an upper bound on the probability that our R.V <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>X</mi></mrow><annotation encoding="application/x-tex">X</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07847em">X</span></span></span></span> spit out a value that was greater than a constant <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>a</mi></mrow><annotation encoding="application/x-tex">a</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em"></span><span class="mord mathnormal">a</span></span></span></span>. This information is useful, but what we really want is the probability that our estimation is between an upper and lower bound. Let's say we have a simple election between candidate A and candidate B. We say that RV <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>X</mi></mrow><annotation encoding="application/x-tex">X</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07847em">X</span></span></span></span> outputs the percentage of the vote candidate A receives. That is, the random variable <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>X</mi></mrow><annotation encoding="application/x-tex">X</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07847em">X</span></span></span></span> outputs a number between 0 and 1. If the actual result of the election is 0.5 (a tie), Markov's inequality may be useful to tell us how likely it is our poll actually tells us that the final vote will 0.5.</p>
<h3>Closing Remarks on the Math</h3>
<p>Markov's and Chebyshev's inequalities are truths. You can of course disagree with the conclusions I draw in this article, and I will even point out places where you should disagree, but the math is unimpeachable. If you have any faith in mathematics, then you must at least accept the definitions above because they all come from the same mathematical foundation. If you have no faith in mathematics, god help you.</p>
<p>Here we begin the process of translating from the impenetrable mathematics to a model we can use to predict and estimate events in the real world. Polling is complicated process expressly because we cannot apply Chebyshev's Inequality perfectly to our world of flesh and blood. The core assumption of Chebyshev's Inequality applied to polling is that we are able to collect a perfectly representative sample of the place we are trying to predict. If we want to predict the U.S. Presidential Election with a single poll, the slice of people we poll needs to perfectly represent the whole country. We need to represent the opinions of an impossibly diverse set of three hundred million people using the voices of only a few. Chebyshev's Inequality, while brilliant, doesn't solve for voting ethno-racial and socioeconomic strata, nor income inequality, aging populations, nor the rural/urban divide.</p>
<h2>The Importance of Counties</h2>
<p>Counties are the smallest voting increment when we talk about the U.S. Presidential election. Each county is responsible for administering the vote within its boundaries. If you live in, say, Cayahoga county, you may only vote in Cayahoga county, and your vote for the presidential election is counted by the county government of Cayahoga county. The county reports the vote total to the state government, and the state government reports the state vote total to the Federal Election Commission. This is true even in the case of absentee ballot voting. If you choose to vote absentee, you send in your ballot to the county where you permanently reside.</p>
<p>Our model relies heavily on U.S. Census data that the Constitution mandates occur every ten years. The U.S. Census data is provided with the finest resolution at the county level, so this a natural limit on the resolution of our model. Even if we wanted to build a model with finer resolution, what standardized unit could we possibly use? The concept of neighborhoods or locals is poorly defined and subject to rapid change. Therefore, we use counties as the finest resolution of our model.tw</p>
<h2>Case Study: The 2016 U.S. Presidential Election</h2>
<p>We're going to perform a thought experiment with the fictitious Middletown County. Middletown county is (self-identified) 70% white voters, 20% Latino voters, and 10% Black voters. Middletown County is home to 1 million people. We are going to make the assumptions that all citizens of Middletown County are eligible to vote and that all citizens of Middletown exercise their right to vote. These assumptions will be addressed later, but for now all they add up to is that <strong>all 1 million citizens of Middletown vote in the 2016 election</strong>. Continuing the thought experiment, let's say we know the end result - Clinton wins the vote in Middletown County with 51.0% of the votes to Trump's 49.0% by a margin of 1.0%. We said before that all of Middletown's citizens vote, therefore Clinton received 510,000 votes, Trump received 490,000 votes, and Clinton won by a margin of 20,000 votes.</p>
<p>Chebyshev's Inequality guarantees that as we poll more people, the probability that our estimation of the final vote percentage is inaccurate, decreases. Let's say it again. The more people we poll, the more likely it is we are going to estimate correctly. The beautiful thing is that we now know how the mathematical tools work, so we can come up with hard numbers and not just say words that sound correct. Let's calculate the exact number of voters we need to poll to all but guarantee that we are below a specific error bound. We decide months before the election that we want to poll the population in a way that guarantees we are accurate to within 2% (total spread), 99% of the time. That means we are within +/- 1% of the actual result with 99% confidence. We both know the actual results of the election, but here are a few more results that would satisfy the above condition of Chebyshev's Inequality.</p>
<ul>
<li>Clinton 50.5% to Trump 49.5% -- Clinton wins -- 1.0% total absolute error</li>
<li>Clinton 50.9% to Trump 49.1% -- Clinton wins -- 0.2% total absolute error</li>
<li>Clinton 50.001% to Trump 49.999% -- Clinton wins -- 1.998% total absolute error</li>
<li>Clinton 50.0% to Trump 49.0% -- Clinton wins -- 0% total absolute error</li>
</ul>
<p>We clearly did an excellent job of predicting the election if made any of the claims above! If our poll came up with any one of the results above, we not only predicted the winner of the election in Middletown County, but we did it with less than 2% absolute error in the spread. Keep in mind the power of the previous statement. If we could actually predict the outcome of elections with this accuracy, I wouldn't be writing this post. I would keep this a secret, start a polling company, and make money predicting the future with a high degree of accuracy. And here's the best part - the number of people we need to poll is a set number. It doesn't matter if the underlying population of Middletown is 1 million, 1 billion, or 1 trillion. There is a specific number of people that all but guarantees (99% is essentially a certainty) that we will predict the result to within +/- 1% accuracy. So what is the magic number?</p>
<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mn>250,000</mn></mrow><annotation encoding="application/x-tex">250{,}000</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8389em;vertical-align:-0.1944em"></span><span class="mord">250</span><span class="mord"><span class="mpunct">,</span></span><span class="mord">000</span></span></span></span></span>
<p>Solving Chebyshev's Inequality for an absolute result error of +/-1% with 99% confidence gives that we need to poll 250,000 people. If we do that, we will estimate the election result to within +/-1%, 99% of the time (equivalent to 2% total spread).</p>
<p>You might be thinking that our work is done. We calculated the number of people we need to poll, and the math checks out. But of course it isn't this easy. The math oversimplifies the numerous and aggravating complexities of the situation. Remember that there are 1 million people in Middletown County. According to the numbers at the beginning, 200,000 of the citizens in Middletown are Latino, and 100,000 are Black. Let's go through a quick sanity check - do we think the Black or Latino communities are going to vote for Donald Trump? Certainly not. There may be a universe where Donald Trump will win the majority of Latino or Black votes, but this is not it. Donald Trump may "love Hispanics!", but they aren't going to turn out to vote for him. Hillary Clinton is essentially guaranteed as a Democrat to win more than 90% of the Black vote. She is also going to win at least 55% of the Latino vote. So if we went to predominately Black and Latino neighborhoods to conduct our poll, and managed to poll 250,000 people to respond, we would deduce that Hillary Clinton is going to win Middletown County in a landslide.</p>
<p>It doesn't require a deep background to see that modern polling rests on assumptions that are nearly impossible to enforce. Below I detail several of the dubious assumptions pollsters make when using traditional polling methodologies. I highly recommend reading the cited literature; each of these assumptions has been comprehensively assessed in other venues.</p>
<h2>Classic Polling Assumptions</h2>
<h3>Voter Turnout</h3>
<p>We decided at the beginning of the case study that every eligible voter in Middletown votes on election day. This is not the case in any real election. Voter turnout has rarely exceeded 60% of eligible voters. For instance, about 55% of eligible voters vote in the U.S. Presidential Election on average. As we move to down-ballot races voter turnout numbers drop dramatically. For local elections voter turnout hovers around single digits. That we naively expected every single person in Middletown County to vote now seems laughable.</p>
<h3>Minority Voter Suppression</h3>
<p>Voter turnout will also be lower than 100% because a raft of deliberate policies and systemic racial issues suppress the votes of minorities. Following the passage of the 13th amendment, a host of Southern states employed literacy tests, poll taxes, and blatant voter intimidation to stifle the African-American vote. Tactics of voter suppression continued through the Jim Crow era until then President LBJ passed the Voting Rights Act of 1964. Among other things, this legislation eliminated all literacy tests and poll taxes, and gave the Justice Department final say on any voting legislation produced in a select list of Southern states with a history of voter suppression tied to systemic racism. In the last thirty years we have seen a concerted behind-the-scenes effort by conservative lawyers to challenge the legality of the Voting Rights Act of 1964, paving the way for the current raft of voter I.D. laws [4]. The tactic of requiring increased identification to vote is a thinly veiled move to suppress minority voters, who are less likely to have the proper identification to satisfy the new criteria. Not coincidentally, voter I.D legislation is cropping up in states with burgeoning minority populations that tend to vote for Democrats. A curious thought is that minority voter suppression engenders a vicious cycle. Republicans at the state and county level push for voter suppression because minorities vote for Democrats by a wide margin, and minorities often vote for Democrats because they are pushing to end minority voter suppression.</p>
<h3>Likely Voters</h3>
<p>Knowing what candidate a citizen would vote for is unhelpful if the citizen doesn't exercise their right to vote. Traditional polls focus heavily on assigning a likelihood to each person surveyed often by simply asking: "How likely are you to vote on or before Election Day?". As an example, we conduct a poll for an upcoming county election with a voting eligible population (VAP) of 50,000. Our 1000 citizen poll results in 100 responses from Black voters, with a margin of <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo>&gt;</mo><mn>90</mn><mi mathvariant="normal">%</mi></mrow><annotation encoding="application/x-tex">&gt;90\%</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.5782em;vertical-align:-0.0391em"></span><span class="mrel">&gt;</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.8056em;vertical-align:-0.0556em"></span><span class="mord">90%</span></span></span></span> saying they would vote for the Democratic candidate. However, none of these poll respondents give unequivocal answers that they will vote on election day. We therefore assign each citizen a low probability of turning out to vote. Viewed in the aggregate, the poll tells us that the majority of Black citizens who vote will vote Democrat, but more importantly that many Black citizens will choose not to vote in our fictitious county election. Turnout among that segment of the population is predicted to be low. We can quickly identify issues with scaling this methodology. In a state with millions of voters, we extrapolate voter turnout based on limited data built on a racial classification system that is anything but nuanced. It would be difficult, say, to defend how the responses of 100 citizens who identify as 'Black' can accurately predict voter turnout of any states' Black population that is both ethnically diverse and geographically spread out.</p>
<h3>Respondent Honesty</h3>
<p>Every poll assumes that the majority of respondents are honest or conversely, that the number of dishonest respondents is non-negligible. Given the choice between Candidate A, Candidate B, and Undecided, poll respondents will provide an honest answer. There is a distinct lack of empirical evidence for this assumption because it is difficult to determine if a respondent's answer to a poll matched his/her actual voting behavior. This assumption deserves a lot of our attention. If we estimate that as low as 0.5% of our respondents will lie, we are already faced with awkward and perplexing choices in our polling methodology. Empirically we have that no racial or ethnic group has any predisposition towards lying that makes that group more likely to lie than any other racial or ethnic group. Therefore, each of our respondents is equally likely to be a liar. The making of a polling catastrophe is quite simple. Say we have a high non-response rate among uneducated White voters in a district where they are a marginal group, yet still represent enough votes to swing an election. We take whatever responses we get from uneducated White voters and extrapolate them to predict the behavior of the voting bloc. But say also that there is a liar among our uneducated White respondents, entirely possible given our uniform distribution of liars among racial and ethnic groups. We have now amplified a skewed result thereby corrupting the accuracy of our poll. If we are unfortunate enough to have lying respondents in our sample of minority groups, we can end up inordinately increasing our overall polling error.</p>
<h2>A Polling Strategy</h2>
<p>The polling establishment is now encountering the perfect storm of all of the issues described in the assumptions above. Response rates are falling making it more expensive to obtain large numbers of valid poll respondents, ethnic and racial groups are increasingly blurred and fragmented making prediction in these categories prone to error, and voter turnout is difficult to assess and impossible to quantify, Traditional polling may still prove effective in small district races, but it is ever more ineffective at the state and national levels. It is both comical and sad to remember that a key argument in favor of modern polling is that all the issues I described above are managed if the probability of respondents lying is low and we have a sizable population of respondents. The results of the most recent election make it safe to say that there were serious issues with the polling in this election cycle.</p>
<p>Hillygus explains succinctly in Public Opinion Quarterly [2]:</p>
<blockquote>
<p>The political parties have built enormous databases that contain information about every registered voter in the United States. Statewide, electronic voter registration files?mandated by the 2002 Help America Vote Act?are the cornerstone of these databases. These files typically include a person's name, home address, turnout history, party registration, phone number, and other information, and are available to parties and candidates (and, in most states, anyone else who wants it). Consumer, census, political, and polling data are then merged into these files to better predict who is going to turn out, what their beliefs and attitudes are and, ultimately, how they are going to vote.</p>
</blockquote>
<p>The old methods of polling are antiquated in the age of big-data and data mining. Rather, we should focus our efforts on building sophisticated statistical tools that have greater predictive power. This is already under way with several major statistical tools being applied to predictive polling. Researchers typically break polls down on based on ethnicity, age, and gender depending on the complexity and targeting of the poll. We assume that there is nothing that differentiates the kind of person who will respond to a poll from someone who won't respond to a poll within ethnic or socioeconomic strata. That is, one White female between the ages of 35-40 must be interchangeable from another woman in the same category. For polling to work, knowing how one woman votes must provide information on how women composing the same group will vote. While it sounds dehumanizing, statistics and history validate that this class of techniques has predictive power. It turns out that we can use Bayesian modeling to figure out how much a specific attribute of an individual contributes to their vote, and then use that to predict of an entire population. We can actually figure out to a precise degree how much being a man/woman, white non-white, old/young, and urban/rural effect your vote likelihood and chosen candidate. By polling thousands of people, I can figure out that in a specific county your likelihood of voting is "20% based on race, 20% based on age, 40% based on gender, and 20% based on a record of government or military service". Our unique characteristics as humans are being used as data-points in a machine that seeks to rob us of the free will we have the right to exercise in an election!</p>
<p>The key idea behind the model we built is that it is incredibly easy to screw up or bias a poll. Polls are very, very difficult execute without introducing error. Our group's model rests on the assumption that polls are highly effective in small geographic areas, and ineffective at predicting larger geographic areas. We contend that wide ranging polls contain too many places to introduce error, and that we can limit our prediction error by executing fewer targeted polls with high accuracy, and then reconstruct the entire voting map with signal estimation techniques. Recognizing that polling requires financial resources, we argue that those resources should be directed at getting highly accurate targeted polls in smaller geographic areas. The question then becomes what counties to select as a representative sample of all American counties. The elegant answer that downplays the difficulty is to find the counties that are the most predictive with respect to historical elections and current economic and demographic data. We built our mathematical model using a network theory approach. Our group has demonstrated that we can more accurately predict the outcome of historical US Presidential elections using highly accurate polls from approximately 50 counties rather than using traditional national polling techniques.</p>
<h2>Notes on Theory</h2>
<p>Refer to Appendix A for the mathematical description of the theory we utilize.</p>
<p>Our model computes multiple data sources to determine how counties voting patterns are correlated. Sources for our data include the CQ Elections database, official FEC Election Reports, and the US Census data. We attempt to only rely on official data to produce our county correlation matrix.</p>
<h2>References</h2>
<p>[1] Kousha Etessami. Markov and Chebyshev's Inequalities.</p>
<p>[2] D. Sunshine Hillygus. "The Evolution of Election Polling in the United States". In: Public Opinion Quarterly 75.5 (2011), pp. 962-981.</p>
<p>[3] Scott Keeter Michael Mokrzycki and Courney Kennedy. "Cell-Phone-Only Voters in the 2008 Exit Poll and Implications for Future Non-coverage Bias". In: Public Opinion Quarterly 73.5 (2009), pp. 845-865.</p>
<p>[4] Jeffrey Toobin. "Holder v. Roberts. The Attorney General Makes Voting Rights the Test Case of his Tenure." In: The New Yorker February (2014).</p>
<h2>Appendix A: Theory</h2>
<p>A brief overview of the core theory and requisite terminology used for this research is presented. Credit to Marques, Segarra, et. al for the presentation of this theory.</p>
<p>A graph <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="script">G</mi></mrow><annotation encoding="application/x-tex">\mathcal{G}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7805em;vertical-align:-0.0972em"></span><span class="mord mathcal" style="margin-right:0.0593em">G</span></span></span></span> is defined as <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="script">G</mi><mo>=</mo><mo stretchy="false">(</mo><mi mathvariant="script">N</mi><mo separator="true">,</mo><mi mathvariant="script">E</mi><mo separator="true">,</mo><mi mathvariant="script">W</mi><mo stretchy="false">)</mo></mrow><annotation encoding="application/x-tex">\mathcal{G} = (\mathcal{N}, \mathcal{E}, \mathcal{W})</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7805em;vertical-align:-0.0972em"></span><span class="mord mathcal" style="margin-right:0.0593em">G</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mopen">(</span><span class="mord mathcal" style="margin-right:0.14736em">N</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord mathcal" style="margin-right:0.08944em">E</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord mathcal" style="margin-right:0.08222em">W</span><span class="mclose">)</span></span></span></span>. The set of nodes <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="script">N</mi></mrow><annotation encoding="application/x-tex">\mathcal{N}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathcal" style="margin-right:0.14736em">N</span></span></span></span> has size <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>N</mi></mrow><annotation encoding="application/x-tex">N</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.10903em">N</span></span></span></span>, the set of edges <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="script">E</mi></mrow><annotation encoding="application/x-tex">\mathcal{E}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathcal" style="margin-right:0.08944em">E</span></span></span></span> is such that edge <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo stretchy="false">(</mo><mi>i</mi><mo separator="true">,</mo><mi>j</mi><mo stretchy="false">)</mo><mo>∈</mo><mi mathvariant="script">E</mi></mrow><annotation encoding="application/x-tex">(i, j) \in \mathcal{E}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mopen">(</span><span class="mord mathnormal">i</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord mathnormal" style="margin-right:0.05724em">j</span><span class="mclose">)</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">∈</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathcal" style="margin-right:0.08944em">E</span></span></span></span> if node <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>i</mi></mrow><annotation encoding="application/x-tex">i</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6595em"></span><span class="mord mathnormal">i</span></span></span></span> is connected to node <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>j</mi></mrow><annotation encoding="application/x-tex">j</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.854em;vertical-align:-0.1944em"></span><span class="mord mathnormal" style="margin-right:0.05724em">j</span></span></span></span>, and every edge in the set <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="script">E</mi></mrow><annotation encoding="application/x-tex">\mathcal{E}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathcal" style="margin-right:0.08944em">E</span></span></span></span> has a corresponding weight in the set <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="script">W</mi></mrow><annotation encoding="application/x-tex">\mathcal{W}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathcal" style="margin-right:0.08222em">W</span></span></span></span>. A signal <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="bold">x</mi><mo>=</mo><msup><mrow><mo fence="true">[</mo><mtable rowspacing="0.16em" columnalign="center center center center" columnspacing="1em"><mtr><mtd><mstyle scriptlevel="0" displaystyle="false"><msub><mi>x</mi><mn>1</mn></msub></mstyle></mtd><mtd><mstyle scriptlevel="0" displaystyle="false"><msub><mi>x</mi><mn>2</mn></msub></mstyle></mtd><mtd><mstyle scriptlevel="0" displaystyle="false"><mo lspace="0em" rspace="0em">⋯</mo></mstyle></mtd><mtd><mstyle scriptlevel="0" displaystyle="false"><msub><mi>x</mi><mi>n</mi></msub></mstyle></mtd></mtr></mtable><mo fence="true">]</mo></mrow><mi>T</mi></msup></mrow><annotation encoding="application/x-tex">\mathbf{x} = \begin{bmatrix} x_1 &amp; x_2 &amp; \cdots &amp; x_n \end{bmatrix}^{T}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4444em"></span><span class="mord mathbf">x</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:1.4312em;vertical-align:-0.35em"></span><span class="minner"><span class="minner"><span class="mopen delimcenter" style="top:0em"><span class="delimsizing size1">[</span></span><span class="mord"><span class="mtable"><span class="col-align-c"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.85em"><span style="top:-3.01em"><span class="pstrut" style="height:3em"></span><span class="mord"><span class="mord"><span class="mord mathnormal">x</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3011em"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight">1</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em"><span></span></span></span></span></span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.35em"><span></span></span></span></span></span><span class="arraycolsep" style="width:0.5em"></span><span class="arraycolsep" style="width:0.5em"></span><span class="col-align-c"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.85em"><span style="top:-3.01em"><span class="pstrut" style="height:3em"></span><span class="mord"><span class="mord"><span class="mord mathnormal">x</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3011em"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight">2</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em"><span></span></span></span></span></span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.35em"><span></span></span></span></span></span><span class="arraycolsep" style="width:0.5em"></span><span class="arraycolsep" style="width:0.5em"></span><span class="col-align-c"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.85em"><span style="top:-3.01em"><span class="pstrut" style="height:3em"></span><span class="mord"><span class="minner">⋯</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.35em"><span></span></span></span></span></span><span class="arraycolsep" style="width:0.5em"></span><span class="arraycolsep" style="width:0.5em"></span><span class="col-align-c"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.85em"><span style="top:-3.01em"><span class="pstrut" style="height:3em"></span><span class="mord"><span class="mord"><span class="mord mathnormal">x</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.1514em"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">n</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em"><span></span></span></span></span></span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.35em"><span></span></span></span></span></span></span></span><span class="mclose delimcenter" style="top:0em"><span class="delimsizing size1">]</span></span></span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:1.0812em"><span style="top:-3.3029em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.13889em">T</span></span></span></span></span></span></span></span></span></span></span></span> is defined on the graph <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="script">G</mi></mrow><annotation encoding="application/x-tex">\mathcal{G}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7805em;vertical-align:-0.0972em"></span><span class="mord mathcal" style="margin-right:0.0593em">G</span></span></span></span> where <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>x</mi><mi>i</mi></msub></mrow><annotation encoding="application/x-tex">x_i</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.5806em;vertical-align:-0.15em"></span><span class="mord"><span class="mord mathnormal">x</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">i</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em"><span></span></span></span></span></span></span></span></span></span> is the value of the signal corresponding to node <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>i</mi></mrow><annotation encoding="application/x-tex">i</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6595em"></span><span class="mord mathnormal">i</span></span></span></span>. Formally, <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="bold">x</mi><mo>∈</mo><msup><mi mathvariant="double-struck">C</mi><mi>N</mi></msup></mrow><annotation encoding="application/x-tex">\mathbf{x} \in \mathbb{C}^N</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.5782em;vertical-align:-0.0391em"></span><span class="mord mathbf">x</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">∈</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.8413em"></span><span class="mord"><span class="mord mathbb">C</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8413em"><span style="top:-3.063em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight" style="margin-right:0.10903em">N</span></span></span></span></span></span></span></span></span></span></span>. We have defined the graph structure, as well as an arbitrary signal defined on the graph.</p>
<p>The graph <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="script">G</mi></mrow><annotation encoding="application/x-tex">\mathcal{G}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7805em;vertical-align:-0.0972em"></span><span class="mord mathcal" style="margin-right:0.0593em">G</span></span></span></span> has a graph-shift-operator <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="bold">S</mi></mrow><annotation encoding="application/x-tex">\mathbf{S}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6861em"></span><span class="mord mathbf">S</span></span></span></span> defined as an <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>S</mi><mo>∈</mo><msup><mi mathvariant="double-struck">R</mi><mrow><mi>N</mi><mo>×</mo><mi>N</mi></mrow></msup></mrow><annotation encoding="application/x-tex">S \in \mathbb{R}^{N \times N}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7224em;vertical-align:-0.0391em"></span><span class="mord mathnormal" style="margin-right:0.05764em">S</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">∈</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.8413em"></span><span class="mord"><span class="mord mathbb">R</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8413em"><span style="top:-3.063em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.10903em">N</span><span class="mbin mtight">×</span><span class="mord mathnormal mtight" style="margin-right:0.10903em">N</span></span></span></span></span></span></span></span></span></span></span></span> matrix satisfying <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>S</mi><mrow><mi>i</mi><mi>j</mi></mrow></msub><mo>=</mo><mn>0</mn></mrow><annotation encoding="application/x-tex">S_{ij} = 0</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.9694em;vertical-align:-0.2861em"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.05764em">S</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em"><span style="top:-2.55em;margin-left:-0.0576em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.05724em">ij</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2861em"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.6444em"></span><span class="mord">0</span></span></span></span> for <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>i</mi><mo mathvariant="normal">≠</mo><mi>j</mi></mrow><annotation encoding="application/x-tex">i \neq j</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em"></span><span class="mord mathnormal">i</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel"><span class="mrel"><span class="mord vbox"><span class="thinbox"><span class="rlap"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em"></span><span class="inner"><span class="mord"><span class="mrel"></span></span></span><span class="fix"></span></span></span></span></span><span class="mrel">=</span></span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.854em;vertical-align:-0.1944em"></span><span class="mord mathnormal" style="margin-right:0.05724em">j</span></span></span></span> and <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo stretchy="false">(</mo><mi>i</mi><mo separator="true">,</mo><mi>j</mi><mo stretchy="false">)</mo><mo mathvariant="normal">∉</mo><mi mathvariant="script">E</mi></mrow><annotation encoding="application/x-tex">(i, j) \notin \mathcal{E}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mopen">(</span><span class="mord mathnormal">i</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord mathnormal" style="margin-right:0.05724em">j</span><span class="mclose">)</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel"><span class="mord"><span class="mrel">∈</span></span><span class="mord vbox"><span class="thinbox"><span class="llap"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="inner"><span class="mord"><span class="mord">/</span><span class="mspace" style="margin-right:0.0556em"></span></span></span><span class="fix"></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathcal" style="margin-right:0.08944em">E</span></span></span></span>. We assume that <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="bold">S</mi></mrow><annotation encoding="application/x-tex">\mathbf{S}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6861em"></span><span class="mord mathbf">S</span></span></span></span> is diagonalizable, so there exists an <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>N</mi><mo>×</mo><mi>N</mi></mrow><annotation encoding="application/x-tex">N \times N</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7667em;vertical-align:-0.0833em"></span><span class="mord mathnormal" style="margin-right:0.10903em">N</span><span class="mspace" style="margin-right:0.2222em"></span><span class="mbin">×</span><span class="mspace" style="margin-right:0.2222em"></span></span><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.10903em">N</span></span></span></span> eigenvector matrix <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="bold">V</mi></mrow><annotation encoding="application/x-tex">\mathbf{V}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6861em"></span><span class="mord mathbf" style="margin-right:0.01597em">V</span></span></span></span> and an <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>N</mi><mo>×</mo><mi>N</mi></mrow><annotation encoding="application/x-tex">N \times N</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7667em;vertical-align:-0.0833em"></span><span class="mord mathnormal" style="margin-right:0.10903em">N</span><span class="mspace" style="margin-right:0.2222em"></span><span class="mbin">×</span><span class="mspace" style="margin-right:0.2222em"></span></span><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.10903em">N</span></span></span></span> eigenvalue matrix <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="bold">Λ</mi></mrow><annotation encoding="application/x-tex">\boldsymbol{\Lambda}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6861em"></span><span class="mord"><span class="mord"><span class="mord mathbf">Λ</span></span></span></span></span></span> that can be used to decompose <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="bold">S</mi></mrow><annotation encoding="application/x-tex">\mathbf{S}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6861em"></span><span class="mord mathbf">S</span></span></span></span> into <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="bold">S</mi><mo>=</mo><mi mathvariant="bold">V</mi><mi mathvariant="bold">Λ</mi><msup><mi mathvariant="bold">V</mi><mrow><mo>−</mo><mn>1</mn></mrow></msup></mrow><annotation encoding="application/x-tex">\mathbf{S} = \mathbf{V}\boldsymbol{\Lambda}\mathbf{V}^{-1}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6861em"></span><span class="mord mathbf">S</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.8141em"></span><span class="mord mathbf" style="margin-right:0.01597em">V</span><span class="mord"><span class="mord"><span class="mord mathbf">Λ</span></span></span><span class="mord"><span class="mord mathbf" style="margin-right:0.01597em">V</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8141em"><span style="top:-3.063em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">−</span><span class="mord mtight">1</span></span></span></span></span></span></span></span></span></span></span></span>. If <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="bold">S</mi></mrow><annotation encoding="application/x-tex">\mathbf{S}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6861em"></span><span class="mord mathbf">S</span></span></span></span> is a normal matrix <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="bold">S</mi><msup><mi mathvariant="bold">S</mi><mi>H</mi></msup><mo>=</mo><msup><mi mathvariant="bold">S</mi><mi>H</mi></msup><mi mathvariant="bold">S</mi></mrow><annotation encoding="application/x-tex">\mathbf{S}\mathbf{S}^{H} = \mathbf{S}^{H}\mathbf{S}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8413em"></span><span class="mord mathbf">S</span><span class="mord"><span class="mord mathbf">S</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8413em"><span style="top:-3.063em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.08125em">H</span></span></span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.8413em"></span><span class="mord"><span class="mord mathbf">S</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8413em"><span style="top:-3.063em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.08125em">H</span></span></span></span></span></span></span></span></span><span class="mord mathbf">S</span></span></span></span>. This means that <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="bold">V</mi></mrow><annotation encoding="application/x-tex">\mathbf{V}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6861em"></span><span class="mord mathbf" style="margin-right:0.01597em">V</span></span></span></span> is unitary and gives <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msup><mi mathvariant="bold">V</mi><mrow><mo>−</mo><mn>1</mn></mrow></msup><mo>=</mo><msup><mi mathvariant="bold">V</mi><mi>H</mi></msup></mrow><annotation encoding="application/x-tex">\mathbf{V}^{-1} = \mathbf{V}^{H}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8141em"></span><span class="mord"><span class="mord mathbf" style="margin-right:0.01597em">V</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8141em"><span style="top:-3.063em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">−</span><span class="mord mtight">1</span></span></span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.8413em"></span><span class="mord"><span class="mord mathbf" style="margin-right:0.01597em">V</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8413em"><span style="top:-3.063em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.08125em">H</span></span></span></span></span></span></span></span></span></span></span></span>. This yields the decomposition of the graph shift operator <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="bold">S</mi><mo>=</mo><mi mathvariant="bold">V</mi><mi mathvariant="bold">Λ</mi><msup><mi mathvariant="bold">V</mi><mi>H</mi></msup></mrow><annotation encoding="application/x-tex">\mathbf{S} = \mathbf{V}\boldsymbol{\Lambda}\mathbf{V}^{H}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6861em"></span><span class="mord mathbf">S</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.8413em"></span><span class="mord mathbf" style="margin-right:0.01597em">V</span><span class="mord"><span class="mord"><span class="mord mathbf">Λ</span></span></span><span class="mord"><span class="mord mathbf" style="margin-right:0.01597em">V</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8413em"><span style="top:-3.063em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.08125em">H</span></span></span></span></span></span></span></span></span></span></span></span>.</p>
<p>The graph shift operator allows for the representation of the signal in the frequency domain, understanding as such, as the use of a basis that is invariant to linear filtering. The Graph Fourier Transform (GFT) is defined as <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mover accent="true"><mi mathvariant="bold">x</mi><mo>~</mo></mover><mo>=</mo><msup><mi mathvariant="bold">V</mi><mrow><mo>−</mo><mn>1</mn></mrow></msup><mi mathvariant="bold">x</mi></mrow><annotation encoding="application/x-tex">\tilde{\mathbf{x}} = \mathbf{V}^{-1}\mathbf{x}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6813em"></span><span class="mord accent"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.6813em"><span style="top:-3em"><span class="pstrut" style="height:3em"></span><span class="mord mathbf">x</span></span><span style="top:-3.3634em"><span class="pstrut" style="height:3em"></span><span class="accent-body" style="left:-0.2222em"><span class="mord">~</span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.8141em"></span><span class="mord"><span class="mord mathbf" style="margin-right:0.01597em">V</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8141em"><span style="top:-3.063em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">−</span><span class="mord mtight">1</span></span></span></span></span></span></span></span></span><span class="mord mathbf">x</span></span></span></span> where <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mover accent="true"><mi mathvariant="bold">x</mi><mo>~</mo></mover></mrow><annotation encoding="application/x-tex">\tilde{\mathbf{x}}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6813em"></span><span class="mord accent"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.6813em"><span style="top:-3em"><span class="pstrut" style="height:3em"></span><span class="mord mathbf">x</span></span><span style="top:-3.3634em"><span class="pstrut" style="height:3em"></span><span class="accent-body" style="left:-0.2222em"><span class="mord">~</span></span></span></span></span></span></span></span></span></span> are the frequency components of the signal. The inverse Graph Fourier Transform (iGFT) is therefore <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="bold">x</mi><mo>=</mo><mi mathvariant="bold">V</mi><mover accent="true"><mi mathvariant="bold">x</mi><mo>~</mo></mover></mrow><annotation encoding="application/x-tex">\mathbf{x} = \mathbf{V}\tilde{\mathbf{x}}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4444em"></span><span class="mord mathbf">x</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.6861em"></span><span class="mord mathbf" style="margin-right:0.01597em">V</span><span class="mord accent"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.6813em"><span style="top:-3em"><span class="pstrut" style="height:3em"></span><span class="mord mathbf">x</span></span><span style="top:-3.3634em"><span class="pstrut" style="height:3em"></span><span class="accent-body" style="left:-0.2222em"><span class="mord">~</span></span></span></span></span></span></span></span></span></span>. We say that the signal <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="bold">x</mi></mrow><annotation encoding="application/x-tex">\mathbf{x}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4444em"></span><span class="mord mathbf">x</span></span></span></span> is <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>K</mi></mrow><annotation encoding="application/x-tex">K</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07153em">K</span></span></span></span>-bandlimited on the graph shift operator <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="bold">S</mi></mrow><annotation encoding="application/x-tex">\mathbf{S}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6861em"></span><span class="mord mathbf">S</span></span></span></span> if its GFT has at most <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>K</mi></mrow><annotation encoding="application/x-tex">K</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07153em">K</span></span></span></span> nonzero components. That is, <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msup><mi mathvariant="bold">V</mi><mrow><mo>−</mo><mn>1</mn></mrow></msup><mi mathvariant="bold">x</mi><mo>=</mo><mover accent="true"><mi mathvariant="bold">x</mi><mo>~</mo></mover><mo>=</mo><msup><mrow><mo fence="true">[</mo><mtable rowspacing="0.16em" columnalign="center center" columnspacing="1em"><mtr><mtd><mstyle scriptlevel="0" displaystyle="false"><msubsup><mover accent="true"><mi mathvariant="bold">x</mi><mo>~</mo></mover><mi>K</mi><mi>T</mi></msubsup></mstyle></mtd><mtd><mstyle scriptlevel="0" displaystyle="false"><msubsup><mn mathvariant="bold">0</mn><mrow><mi>N</mi><mo>−</mo><mi>K</mi></mrow><mi>T</mi></msubsup></mstyle></mtd></mtr></mtable><mo fence="true">]</mo></mrow><mi>T</mi></msup></mrow><annotation encoding="application/x-tex">\mathbf{V}^{-1}\mathbf{x} = \tilde{\mathbf{x}} = \begin{bmatrix} \tilde{\mathbf{x}}_{K}^{T} &amp; \mathbf{0}_{N-K}^{T} \end{bmatrix}^{T}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8141em"></span><span class="mord"><span class="mord mathbf" style="margin-right:0.01597em">V</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8141em"><span style="top:-3.063em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">−</span><span class="mord mtight">1</span></span></span></span></span></span></span></span></span><span class="mord mathbf">x</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.6813em"></span><span class="mord accent"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.6813em"><span style="top:-3em"><span class="pstrut" style="height:3em"></span><span class="mord mathbf">x</span></span><span style="top:-3.3634em"><span class="pstrut" style="height:3em"></span><span class="accent-body" style="left:-0.2222em"><span class="mord">~</span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:1.4326em;vertical-align:-0.3507em"></span><span class="minner"><span class="minner"><span class="mopen delimcenter" style="top:0em"><span class="delimsizing size1">[</span></span><span class="mord"><span class="mtable"><span class="col-align-c"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.8507em"><span style="top:-3.0093em"><span class="pstrut" style="height:3em"></span><span class="mord"><span class="mord"><span class="mord accent"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.6813em"><span style="top:-3em"><span class="pstrut" style="height:3em"></span><span class="mord mathbf">x</span></span><span style="top:-3.3634em"><span class="pstrut" style="height:3em"></span><span class="accent-body" style="left:-0.2222em"><span class="mord">~</span></span></span></span></span></span></span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.8413em"><span style="top:-2.4247em;margin-left:0em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.07153em">K</span></span></span></span><span style="top:-3.063em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.13889em">T</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2753em"><span></span></span></span></span></span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.3507em"><span></span></span></span></span></span><span class="arraycolsep" style="width:0.5em"></span><span class="arraycolsep" style="width:0.5em"></span><span class="col-align-c"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.8507em"><span style="top:-3.0093em"><span class="pstrut" style="height:3em"></span><span class="mord"><span class="mord"><span class="mord mathbf">0</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.8413em"><span style="top:-2.4247em;margin-left:0em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.10903em">N</span><span class="mbin mtight">−</span><span class="mord mathnormal mtight" style="margin-right:0.07153em">K</span></span></span></span><span style="top:-3.063em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.13889em">T</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.3337em"><span></span></span></span></span></span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.3507em"><span></span></span></span></span></span></span></span><span class="mclose delimcenter" style="top:0em"><span class="delimsizing size1">]</span></span></span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:1.0819em"><span style="top:-3.3036em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.13889em">T</span></span></span></span></span></span></span></span></span></span></span></span> where <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mn mathvariant="bold">0</mn><mrow><mi>N</mi><mo>−</mo><mi>K</mi></mrow></msub></mrow><annotation encoding="application/x-tex">\mathbf{0}_{N-K}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8528em;vertical-align:-0.2083em"></span><span class="mord"><span class="mord mathbf">0</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3283em"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.10903em">N</span><span class="mbin mtight">−</span><span class="mord mathnormal mtight" style="margin-right:0.07153em">K</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2083em"><span></span></span></span></span></span></span></span></span></span> is a vector of <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>N</mi><mo>−</mo><mi>K</mi></mrow><annotation encoding="application/x-tex">N - K</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7667em;vertical-align:-0.0833em"></span><span class="mord mathnormal" style="margin-right:0.10903em">N</span><span class="mspace" style="margin-right:0.2222em"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em"></span></span><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07153em">K</span></span></span></span> zeros. If this is the case then the original signal can be reconstructed from a sampled version of the complete signal. Define a <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>K</mi><mo>×</mo><mi>N</mi></mrow><annotation encoding="application/x-tex">K \times N</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7667em;vertical-align:-0.0833em"></span><span class="mord mathnormal" style="margin-right:0.07153em">K</span><span class="mspace" style="margin-right:0.2222em"></span><span class="mbin">×</span><span class="mspace" style="margin-right:0.2222em"></span></span><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.10903em">N</span></span></span></span> selection matrix <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="bold">C</mi></mrow><annotation encoding="application/x-tex">\mathbf{C}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6861em"></span><span class="mord mathbf">C</span></span></span></span> that samples <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>K</mi></mrow><annotation encoding="application/x-tex">K</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07153em">K</span></span></span></span> of <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>N</mi></mrow><annotation encoding="application/x-tex">N</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.10903em">N</span></span></span></span> of the graph nodes. Observe that <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="bold">C</mi></mrow><annotation encoding="application/x-tex">\mathbf{C}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6861em"></span><span class="mord mathbf">C</span></span></span></span> is a selection matrix if it is binary, has exactly one nonzero entry per row and at most one nonzero entry per column. This sampled signal is defined as <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mover accent="true"><mi mathvariant="bold">x</mi><mo>ˉ</mo></mover><mo>=</mo><mi mathvariant="bold">C</mi><mi mathvariant="bold">x</mi></mrow><annotation encoding="application/x-tex">\bar{\mathbf{x}} = \mathbf{C}\mathbf{x}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.5812em"></span><span class="mord accent"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.5812em"><span style="top:-3em"><span class="pstrut" style="height:3em"></span><span class="mord mathbf">x</span></span><span style="top:-3.0134em"><span class="pstrut" style="height:3em"></span><span class="accent-body" style="left:-0.2222em"><span class="mord">ˉ</span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.6861em"></span><span class="mord mathbf">Cx</span></span></span></span>. The reconstruction can be carried out by</p>
<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mi mathvariant="bold">x</mi><mo>=</mo><msub><mi mathvariant="bold">V</mi><mi>K</mi></msub><msub><mover accent="true"><mi mathvariant="bold">x</mi><mo>~</mo></mover><mi>K</mi></msub><mo>=</mo><msub><mi mathvariant="bold">V</mi><mi>K</mi></msub><mo stretchy="false">(</mo><mi mathvariant="bold">C</mi><msub><mi mathvariant="bold">V</mi><mi>K</mi></msub><msup><mo stretchy="false">)</mo><mrow><mo>−</mo><mn>1</mn></mrow></msup><mover accent="true"><mi mathvariant="bold">x</mi><mo>~</mo></mover></mrow><annotation encoding="application/x-tex">\mathbf{x} = \mathbf{V}_{K}\tilde{\mathbf{x}}_{K} = \mathbf{V}_{K}(\mathbf{C}\mathbf{V}_{K})^{-1}\tilde{\mathbf{x}}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4444em"></span><span class="mord mathbf">x</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.8361em;vertical-align:-0.15em"></span><span class="mord"><span class="mord mathbf" style="margin-right:0.01597em">V</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3283em"><span style="top:-2.55em;margin-left:-0.016em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.07153em">K</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em"><span></span></span></span></span></span></span><span class="mord"><span class="mord accent"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.6813em"><span style="top:-3em"><span class="pstrut" style="height:3em"></span><span class="mord mathbf">x</span></span><span style="top:-3.3634em"><span class="pstrut" style="height:3em"></span><span class="accent-body" style="left:-0.2222em"><span class="mord">~</span></span></span></span></span></span></span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3283em"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.07153em">K</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:1.1141em;vertical-align:-0.25em"></span><span class="mord"><span class="mord mathbf" style="margin-right:0.01597em">V</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3283em"><span style="top:-2.55em;margin-left:-0.016em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.07153em">K</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em"><span></span></span></span></span></span></span><span class="mopen">(</span><span class="mord mathbf">C</span><span class="mord"><span class="mord mathbf" style="margin-right:0.01597em">V</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3283em"><span style="top:-2.55em;margin-left:-0.016em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.07153em">K</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em"><span></span></span></span></span></span></span><span class="mclose"><span class="mclose">)</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8641em"><span style="top:-3.113em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">−</span><span class="mord mtight">1</span></span></span></span></span></span></span></span></span><span class="mord accent"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.6813em"><span style="top:-3em"><span class="pstrut" style="height:3em"></span><span class="mord mathbf">x</span></span><span style="top:-3.3634em"><span class="pstrut" style="height:3em"></span><span class="accent-body" style="left:-0.2222em"><span class="mord">~</span></span></span></span></span></span></span></span></span></span></span>
<p>whenever a selection matrix <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="bold">C</mi></mrow><annotation encoding="application/x-tex">\mathbf{C}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6861em"></span><span class="mord mathbf">C</span></span></span></span> that satisfies <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="normal">rank</mi><mo>⁡</mo><mo stretchy="false">{</mo><mi mathvariant="bold">C</mi><msub><mi mathvariant="bold">V</mi><mi>K</mi></msub><mo stretchy="false">}</mo><mo>=</mo><mi>K</mi></mrow><annotation encoding="application/x-tex">\operatorname{rank}\{\mathbf{C}\mathbf{V}_{K}\} = K</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mop"><span class="mord mathrm">rank</span></span><span class="mopen">{</span><span class="mord mathbf">C</span><span class="mord"><span class="mord mathbf" style="margin-right:0.01597em">V</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3283em"><span style="top:-2.55em;margin-left:-0.016em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.07153em">K</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em"><span></span></span></span></span></span></span><span class="mclose">}</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathnormal" style="margin-right:0.07153em">K</span></span></span></span> is used.</p>
<p>The difficulties specific to this theory include generating a selection sampling matrix <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="bold">C</mi></mrow><annotation encoding="application/x-tex">\mathbf{C}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6861em"></span><span class="mord mathbf">C</span></span></span></span> that satisfies the above equation, working with a graph signal that is not cleanly bandlimited, and working with a graph shift operator that does not decompose to a full rank eigenvalue matrix.</p>]]></content:encoded>
            <author>gmgprivacy@proton.me (Matt Goodman)</author>
        </item>
        <item>
            <title><![CDATA[What Matters?]]></title>
            <link>https://goodmattg.xyz/articles/What-Matters</link>
            <guid>https://goodmattg.xyz/articles/What-Matters</guid>
            <pubDate>Sat, 20 Dec 2025 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<blockquote>
<p>The unexamined life is not worth living. - Socrates</p>
</blockquote>
<p>This might make me the worst / most-SF person at every party, but my favorite question a few drinks in is to ask “what do you care about?”. The right time to ask is after two beers - first beer is to loosen everyone up, second is to get the honest answers. The first look I get is confused. People usually tilt their heads as if to say “Matt - we’re at a party that’s a stupid question”, god, family, football. But then I’ll press. Beyond your family and community and the obvious things that give life meaning, what matters to you, what do you believe? I think what shakes me is how few of us can put words to our beliefs or to things that we anchor to. In one way I get it - most people are entirely focused on the core things, because that’s how life falls. If you’re raising a kid, the kid is what matters - America’s competitiveness in rare earths matters too - but does it really? The thing I love about my friends is that to some degree everyone has sat down and come up with an answer - it may be a stupid and inane answer, it might just be a grift, but it’s an answer.</p>]]></content:encoded>
            <author>gmgprivacy@proton.me (Matt Goodman)</author>
        </item>
    </channel>
</rss>