<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en_US"><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://blog.sheetsolved.com/feed.xml" rel="self" type="application/atom+xml" /><link href="https://blog.sheetsolved.com/" rel="alternate" type="text/html" hreflang="en_US" /><updated>2026-09-18T11:41:56+00:00</updated><id>https://blog.sheetsolved.com/feed.xml</id><title type="html">Tech Perspectives</title><subtitle>Ces&apos;s random twalks on tech and stats.</subtitle><author><name>Cesaire Tobias</name></author><entry><title type="html">In the Long Run: Which Doomsday Warnings Are Worth Having</title><link href="https://blog.sheetsolved.com/doomsday-warnings.html" rel="alternate" type="text/html" title="In the Long Run: Which Doomsday Warnings Are Worth Having" /><published>2026-09-18T00:00:00+00:00</published><updated>2026-09-18T00:00:00+00:00</updated><id>https://blog.sheetsolved.com/doomsday-warnings</id><content type="html" xml:base="https://blog.sheetsolved.com/doomsday-warnings.html"><![CDATA[<p>The steady stream of headlines about AI bringing about humanity’s demise has stirred up some faint recollection of my economics training - in particular, John Maynard Keynes’s “In the long run we are all dead.” I hear the misquote groans already, but if you give me a chance, I’ve developed this thought a bit further below.</p>

<p>We need people sounding alarms, because being aware of what could go wrong is how we stop it happening. The warnings that do the most good share two things: they say how the disaster would happen, and they say what could be done about it. That gives people something to act on, and when they act, the warning often ends up looking wrong.</p>

<p>Economic historians, researchers who study emergency warnings, forecasters and AI safety researchers have each worked out a piece of it. So, like in most of my writing, I attempt here to synthesise those thoughts.</p>

<h2 id="the-malthusian-trap">The Malthusian trap</h2>

<p>In 1798 Thomas Malthus published <a href="https://www.gutenberg.org/ebooks/4239"><em>An Essay on the Principle of Population</em></a>. His argument was that population multiplies (growing exponentially) while food supply only grows by adding a bit each year (linear growth), so people always end up outnumbering the food until famine, disease or war brings the numbers back down. Judged against the history he had to go on, that was a reasonable reading.</p>

<p>What he didn’t allow for was technology changing how fast food production could grow. He treated the limit as fixed, which, with the benefit of hindsight, we know it wasn’t. So, in assessing any doomsday forecast, the lesson learned is to ask whether it takes today’s limits and projects them forward as though nobody will respond to them.</p>

<h2 id="a-warning-rather-than-a-prophecy">A warning rather than a prophecy</h2>

<p>A hundred years after Malthus, the chemist William Crookes used a major scientific address in 1898 to warn that England and “all civilised nations stand in deadly peril of not having enough to eat.” Wheat needs nitrogen, the Chilean mineral deposits used as fertiliser were running out, and on his figures wheat would fall behind population after 1931. He also said what would fix it: “The fixation of atmospheric nitrogen therefore is one of the great discoveries awaiting the ingenuity of chemists” (<a href="https://archive.org/details/wheatproblembase00croouoft"><em>The Wheat Problem</em></a>, the book version of the address). In other words, find a way to turn the nitrogen in the air into fertiliser.</p>

<p>That’s what happened. By 1913 the German chemical company BASF was running Fritz Haber and Carl Bosch’s process for doing exactly that at industrial scale. Though, I can’t say with certainty that Crookes’s speech prompted their work. Haber’s <a href="https://www.nobelprize.org/uploads/2018/06/haber-lecture.pdf">Nobel lecture</a> describes the same worry about the Chilean deposits becoming “clearly apparent at the turn of the century”, without naming Crookes. Either way, Crookes named the fix that came.</p>

<p>Jason Crawford tells this story in <a href="https://newsletter.rootsofprogress.org/p/solutionism-part-1"><em>Solutionism</em></a>. He points out that the 1931 shortfall was avoided mostly by things Crookes didn’t foresee: tractors made it pay to farm more land, and new wheat varieties helped. He then sets Crookes against Paul Ehrlich’s 1968 bestseller <em>The Population Bomb</em>. In the passages Crawford quotes, the book opens by declaring “the battle to feed all of humanity is over” and later endorses forcing population control on people: “Coercion? Perhaps, but coercion in a good cause.” Crawford’s distinction is that Crookes’s alarm was contingent, meaning it depended on facts that could change, so “when the facts changed, the alarm could end”. Ehrlich’s, in Crawford’s reading, was “impervious to facts”. Crookes described his own speech, in Crawford’s quotation of him, as taking “the form of a warning rather than of a prophecy”: a warning comes with a way to stop it coming true.</p>

<h2 id="what-the-research-says-about-warnings">What the research says about warnings</h2>

<p>Researchers who study emergency warnings for floods, storms and evacuations reached a similar conclusion from a different direction. A <a href="https://www.nationalacademies.org/read/24935/chapter/3">2018 National Academies review</a> of that work says warnings are more likely to get people to act if they cover “guidance, time, location, hazard and consequences, and source”, and that “people increase responsiveness when they receive guidance on exactly what to do.”</p>

<p>Health researchers have studied fear directly. Scaring people does work, if modestly: a <a href="https://pubmed.ncbi.nlm.nih.gov/26501228/">2015 review of 127 studies</a> found fear-based messages generally changed attitudes and behaviour, and found no case where they backfired. But they worked better when they told people what they could do. An <a href="https://doi.org/10.1177/109019810002700506">earlier review</a> found that strong fear plus a clear, workable action produced the most change, while strong fear with no workable action produced the most defensive reactions, such as avoiding the message altogether.</p>

<p>On that evidence, a vague warning does less than it could.</p>

<p>From forecasting, the political scientist Philip Tetlock spent years scoring expert predictions, and <a href="https://www.edge.org/conversation/philip_tetlock-a-short-course-in-superforecasting">found</a> that pundits protect themselves “by relying on vague verbiage. They can often be wrong, but never in error.” A warning that never says what would happen, or when, can’t be checked, so nobody can learn from it.</p>

<h2 id="wrong-because-they-were-heeded">Wrong because they were heeded</h2>

<p>The ozone layer is the clearest case of a specific warning working. In 1974 the chemists Mario Molina and Sherwood Rowland <a href="https://doi.org/10.1038/249810a0">showed</a> that CFCs, gases then used in aerosol cans and fridges, could destroy the ozone that shields the Earth from ultraviolet light. That named a mechanism and a fix: stop using CFCs. The warning didn’t act alone. The <a href="https://doi.org/10.1038/315207a0">ozone hole</a> over Antarctica, damage people could actually see, was reported in 1985, two years before the <a href="https://ozone.unep.org/treaties/montreal-protocol">Montreal Protocol</a> to phase the gases out was signed. The UN’s 2023 assessment says the ozone layer is <a href="https://www.unep.org/news-and-stories/press-release/ozone-layer-recovery-track-helping-avoid-global-warming-05degc">on track to recover</a> within decades.</p>

<p>Looking back, it’s easy to see that as a scare that came to nothing, when the warning did its job. The sociologist Robert Merton had a name for this in 1936, the “suicidal prophecy”, now usually called a <a href="https://users.ox.ac.uk/~sfos0060/papers/prophecies_text.pdf">self-defeating prophecy</a>: a prediction that changes behaviour enough to make itself untrue.</p>

<p>That’s why “doom predictions have always been wrong” is a weak argument. Some were wrong because the danger wasn’t real and some because the warning worked, and the record doesn’t say which is which. Even the famous bet between Ehrlich and the economist Julian Simon over whether resource prices would rise, which Ehrlich lost, proves less than it seems: <a href="https://ourworldindata.org/simon-ehrlich-bet">Our World in Data</a> found that over other decades since 1900 each man would have won about half the time.</p>

<h2 id="two-kinds-of-ai-warning">Two kinds of AI warning</h2>

<p>It’s easy to assume most AI doom names no mechanism at all. These are the clickbait headlines most of us come across each day. That turns out to be unfair to the researchers. AI safety has its own push for concreteness, going back at least to a 2016 paper literally titled <a href="https://arxiv.org/abs/1606.06565"><em>Concrete Problems in AI Safety</em></a>. A <a href="https://arxiv.org/abs/2306.12001">2023 overview</a> sorts catastrophic AI risk into four groups: people deliberately misusing AI, companies and countries racing to deploy it before it’s safe, accidents inside the organisations building it, and AI systems that can’t be controlled. For each it describes specific hazards and proposes practical fixes.</p>

<p>What reaches most people is a different register. A widely reported statement on AI risk, <a href="https://aistatement.com/">signed in 2023</a> by Geoffrey Hinton, Yoshua Bengio and the heads of OpenAI and Google DeepMind among many others, is one sentence long: “Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war.” It names no mechanism and no action, and it’s the version I see in the headlines far more often than anything from the papers.</p>

<p>The best case I found for staying vague comes from Eliezer Yudkowsky, one of the most prominent voices on AI risk. In a <a href="https://time.com/6266923/ai-eliezer-yudkowsky-open-letter-not-enough/">2023 piece for TIME</a> he compares humanity facing a superhuman AI to “a 10-year-old trying to play chess against Stockfish 15”, a chess program far stronger than any human. You can be sure you’ll lose without being able to say which moves will beat you, and knocking down one predicted path to disaster wouldn’t make the danger go away. That’s a fair point. His proposed actions are anything but vague, though: “Shut down all the large GPU clusters.” When you can’t predict the path, the only lever left is whether to start the game at all.</p>

<p>Specificity has a trap of its own as well. Amos Tversky and Daniel Kahneman <a href="https://doi.org/10.1037/0033-295X.90.4.293">showed in 1983</a> that adding detail to a scenario makes it feel more likely even as it becomes less likely, because every extra detail is one more thing that has to come true. A vivid story of exactly how AI goes wrong is more persuasive and less probable than a general one. So specificity says nothing about whether a warning is right. What it adds is that the warning can be checked and acted on.</p>

<p>A smaller example of the useful kind: “AI agents that act on content they read can be steered by instructions hidden in that content.” Left alone, that leads to AI assistants leaking data and following attackers’ orders. It names the mechanism, and people are building defences against it right now. If those defences work, the warning will look wrong in a few years, just as Crookes’s does.</p>

<h2 id="keynes-read-properly">Keynes, read properly</h2>

<p>See, I’ve come back to Keynes. The passage in <em>A Tract on Monetary Reform</em> (1923) reads “But this long run is a misleading guide to current affairs. In the long run we are all dead.” He was writing about economists who said the quantity theory of money (the idea that prices rise in step with the amount of money in circulation) would hold in the long run, and left it there. His target was people using the long run to shrug off present problems. A doomsayer makes the opposite claim: that the long run is catastrophic. The first half of his quote applies to both, though. A long-run forecast, optimistic or apocalyptic, is a poor guide to what to do now.</p>

<h2 id="a-warning-worth-having-about-ai">A warning worth having about AI</h2>

<p>Writing <a href="running-to-stand-still.html"><em>Running to Stand Still</em></a> left me with one I’d put forward. AI is adding code faster than our ways of checking it have adapted, the code that gets through unchecked carries failures and security holes, and the way out is to make checking much cheaper or to need less code.</p>

<p>The evidence for the first part is real, if indirect. At Google, executives say the share of new code written by AI went from <a href="https://blog.google/inside-google/message-ceo/alphabet-earnings-q3-2024/">“more than a quarter”</a> in October 2024 to <a href="https://s206.q4cdn.com/479360582/files/doc_events/2026/Feb/04/2025_Q4_Earnings_Transcript.pdf">about half</a> by early 2026, with engineers reviewing it. Google’s own DevOps research team found in its <a href="https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report">2025 survey</a> of technology professionals that AI adoption goes with less stable software releases, and put that down to checking: without strong automated testing and fast feedback, “an increase in change volume leads to instability.” A <a href="https://arxiv.org/abs/2511.04427">study of open-source projects</a> that adopted the AI coding tool Cursor found a large but short-lived jump in how fast they worked, a lasting rise in code-quality warnings, and named “quality assurance as a major bottleneck.” On the second part, a <a href="https://arxiv.org/abs/2211.03622">Stanford user study</a> found that people with an AI assistant wrote less secure code than people without one, and were more likely to believe their code was secure.</p>

<p>The same DevOps research also points the other way. Its <a href="https://cloud.google.com/blog/products/devops-sre/announcing-the-2024-dora-report">2024 report</a> estimated that as AI adoption rose, code reviews got a few per cent faster and code quality a few per cent better, even while releases got less stable. By 2025, AI adoption also went with teams delivering software faster, which the researchers read as people “learning where, when, and how AI is most useful.” Only the link to less stable releases held. So the evidence shows checking under strain, and some sign of it catching up. It doesn’t show checking falling further behind.</p>

<p>Left alone, unchecked code piles up, and so do the failures and security holes hiding in it. That’s a Malthusian trap, with writing code in the role of population and checking it in the role of food.</p>

<p>Malthus missed two things, and each has a counterpart here. The first was a jump in how fast the limited resource could grow: fertiliser changed food production. For code, the equivalent would be checking getting much cheaper, most plausibly through machines writing mathematical proofs that their code is correct, which a small, simple program can then confirm. The second was demand: once people got richer they chose to have fewer children, which Malthus never expected. For code, the equivalent is needing less of it, where more requests get answered with a finished result rather than a program somebody has to trust.</p>

<p>By this post’s own test, the warning falls short of Crookes’s in one way: it has no date. The evidence shows the strain already appearing, and says nothing about when, or whether, it turns into a crisis. I can’t tell you which way out will come either, or whether it’ll be something that isn’t on anyone’s list yet, just as Malthus had no way to foresee synthetic fertiliser.</p>]]></content><author><name>Cesaire Tobias</name></author><category term="ai" /><category term="economics" /><category term="forecasting" /><summary type="html"><![CDATA[Warnings about disaster do the most good when they say how it would happen and what to do about it. What economic history, warning research, forecasting and AI safety know about that.]]></summary></entry><entry><title type="html">Running to Stand Still: Why Perfect Code Still Gets Patched</title><link href="https://blog.sheetsolved.com/running-to-stand-still.html" rel="alternate" type="text/html" title="Running to Stand Still: Why Perfect Code Still Gets Patched" /><published>2026-09-18T00:00:00+00:00</published><updated>2026-09-18T00:00:00+00:00</updated><id>https://blog.sheetsolved.com/running-to-stand-still</id><content type="html" xml:base="https://blog.sheetsolved.com/running-to-stand-still.html"><![CDATA[<p>It seems that a growing share of code is now written by machines rather than people, and it keeps getting better. That got me wondering whether it could get good enough that software updates would only ever add features. Perfect code would have no bugs to fix, and no vulnerabilities to patch either. But could such a thing as perfect code ever exist?</p>

<p>I’ve concluded not, but for a reason other than I initially expected. Code can get very close to perfect, as long as perfect means doing exactly what it was meant to do at the time it was written. Updates still won’t be features only, because some changes you choose and others the world forces on you. What AI may change is who handles the forced ones. They could turn into something closer to an immune response, happening in the background without anyone deciding on it.</p>

<p>Most of the ideas here come from people who work much closer to these fields than I do, and I’ve credited them as they come up. What I’ve tried to do is put them together in a way that’s digestible for a broad audience, who, like me, are interested but not experts.</p>

<h2 id="perfect-against-what">Perfect against what</h2>

<p>Code can only be correct against something: a specification (spec), which is a precise description of what it should do. Proving code correct against its spec is already possible for some software. seL4, the core of an operating system, and CompCert, a compiler (the program that turns the code people write into instructions a computer runs), both come with mathematical proofs that the code does what its spec says. When researchers fed randomly generated programs to a range of C compilers to hunt for bugs, CompCert was <a href="https://users.cs.utah.edu/~regehr/papers/pldi11-preprint.pdf">the only one</a> where they couldn’t make the proven part produce wrong results. The bugs they did find were in a part the proof didn’t cover, and in a corner the spec didn’t describe fully.</p>

<p>Those proofs hold under stated assumptions, though. The seL4 team <a href="https://www.sigops.org/s/conferences/sosp/2009/papers/klein-sosp09.pdf">lists theirs</a>: the compiler used to build it, the start-up code and the hardware itself are among the things assumed to work correctly. And the spec has to say what you actually wanted. So a proof moves the question to the spec.</p>

<p>It’s tempting to think that’s where perfect code breaks down, because specs come from people and people get things wrong. Say I build an app to sell hotdogs and it sells hotdogs. Later I realise I should have been selling burgers. Nothing was wrong with the hotdog app. I changed my mind, and the burger version is a new release with that burger-selling feature.</p>

<p>What does count against the code is intent I had at the time and never wrote down. If the hotdog app lets someone order minus three hotdogs and collect a refund, I didn’t change my mind about that. I never wanted it, and the spec never said so. A lot of security vulnerabilities look like this: specs describe what should happen, and rarely list everything that must not.</p>

<p>Machines can close most of that gap, because most unstated intent is shared: no negative quantities, no charging twice, no showing one customer another’s orders. In my experience AI models are good at inferring those, and at finding the gaps and contradictions in a spec. It’s why I now write a scope file before starting a project (see <a href="bounding-ai-code-reviews.html"><em>How Long Is a Piece of String?</em></a>). With a machine checking the spec and a proof tying the code to it, code that is correct against what you intended at the time looks achievable. It rests on proof tools getting much cheaper, which I think we can foresee happening over the coming years.</p>

<h2 id="decided-or-imposed">Decided or imposed</h2>

<p>Even if you grant all of that, updates don’t become features only. MD5 and SHA-1 are hash algorithms, used to check that files and digital signatures haven’t been tampered with. A perfectly correct implementation of either is still insecure today, because both were broken after they shipped: MD5 <a href="https://eprint.iacr.org/2004/199">in 2004</a> and SHA-1 <a href="https://shattered.io/">in 2017</a>. RSA, the encryption behind much of the internet’s security, is <a href="https://nvlpubs.nist.gov/nistpubs/ir/2016/NIST.IR.8105.pdf">expected to follow</a> once quantum computers get large enough, which is why the US standards body NIST has already <a href="https://www.nist.gov/news-events/news/2024/08/nist-releases-first-3-finalized-post-quantum-encryption-standards">published replacement standards</a>. None of that code had a bug. The world it runs in changed.</p>

<p>You could argue that “I now want to be safe from the new attack” is just another change of intent, which would make it a feature too. Push that far, though, and every change is a feature, so perfect code is perfect by definition and the question stops meaning anything. Perhaps a more useful framing of change:</p>

<ul>
  <li><strong>Decided:</strong> I chose to change it. Burgers instead of hotdogs.</li>
  <li><strong>Imposed:</strong> the change is forced on me, and I’d lose something by not making it. An algorithm gets broken, someone else’s code that mine relies on gets withdrawn, a regulator changes the rules.</li>
</ul>

<p>This framing is at least fifty years old. E. Burton Swanson <a href="http://www.mit.jyu.fi/ope/kurssit/TIES462/Materiaalit/Swanson.pdf">split software maintenance</a> in 1976 into corrective work (fixing faults), adaptive work (keeping up with a changing environment) and perfective work (improvements people ask for). He described failures and environmental change as causes where “a response is typically unavoidable”, against changes that reflect “the initiatives of user and maintenance personnel”. Manny Lehman’s <a href="https://users.ece.utexas.edu/~perry/education/SE-Intro/lehman.pdf">first law of software evolution</a> says much the same: a program used in the real world “undergoes continual change or becomes progressively less useful.”</p>

<p>What perfect code does to Swanson’s categories is take corrective work, fixing bugs, to near zero. It does nothing about adaptive work, which arrives at the speed the world moves, however good the code is.</p>

<h2 id="adapt-or-perish">Adapt or perish</h2>

<p>Put that way, software starts to look like biology. The environment changes and an organism adapts or dies out. Treating software as an evolving ecosystem has a research literature of its own, and a <a href="https://arxiv.org/abs/2512.02953">recent example</a> asks what AI coding tools will do to it. Security is a close match for what biologists call the Red Queen hypothesis, after the character in <em>Through the Looking-Glass</em> who has to keep running to stay in the same place. Attackers and defenders evolve against each other, and patching is the running. I’m not the first to <a href="https://learn.microsoft.com/en-us/archive/blogs/tzink/the-red-queen-theory-of-internet-security">make that comparison</a> either.</p>

<p>Biology has no perfect organism, only organisms well suited to the environment they’re in right now. Perfect code works the same way. It can be perfect against its spec at a point in time, and it stops being perfect when the world moves.</p>

<p>The analogy breaks on intent. Evolution has none: variation is random and selection does the work. Software has my “decided” changes too, and they get written straight into the next version. That makes software closer to Lamarck, who thought creatures passed on traits they acquired during their lives, than to Darwin.</p>

<h2 id="an-immune-system-for-code">An immune system for code</h2>

<p>Until now, both kinds of change went through a person. Someone decided on the burgers, and someone noticed the broken hash algorithm and wrote the patch. With AI writing the code, the imposed changes don’t need a person any more. A security hole is announced, the affected code gets updated, the automated tests pass and the fix ships. That works more like an immune system than evolution - your body fights off a cold without you deciding to.</p>

<p>Stephanie Forrest and colleagues were building security on immune-system principles in the 1990s (<a href="https://doi.org/10.1145/262793.262811"><em>Computer Immunology</em></a>, 1997), and software that repairs bugs by itself has been a research field for over fifteen years. It now runs end to end without a person. In DARPA’s AI Cyber Challenge, which <a href="https://www.darpa.mil/news/2025/aixcc-results">finished in August 2025</a>, the finalists’ autonomous systems between them found 18 real vulnerabilities in open-source software and supplied patches for 11.</p>

<p>So from a person’s point of view, updates may well become features only. The imposed changes still happen, but as background maintenance nobody decides on or sees.</p>

<h2 id="when-the-immune-system-is-the-target">When the immune system is the target</h2>

<p>An immune system that runs on its own, while convenient, comes with its own perils and becomes a surface worth attacking. Whatever comes in through the update pipeline gets trusted and spread everywhere, so that’s where an attacker should aim. The xz utils backdoor, found in 2024, was an early example: a contributor spent <a href="https://research.swtch.com/xz-timeline">over two years</a> earning the trust of the people who maintain xz, a compression tool built into many Linux systems, then slipped a backdoor in through its normal release process. Automated repair has its own version. A 2025 study <a href="https://arxiv.org/abs/2509.05372">wrote 51 fake bug reports</a> aimed at an AI repair system, and 90% of them got it to produce the patch the attacker wanted. In biology the equivalent is a pathogen that exploits the immune response, or an autoimmune disease where the defences attack the body.</p>

<p>The automation also speeds up the environment it’s responding to. When each system adapts on its own, its changes become imposed changes for everything that depends on it, and those systems adapt in turn. The faster systems adapt, the faster the environment around them changes.</p>

<p>Having the same models write the spec, the code and the review makes this worse. As I found with AI code reviews, each run is one sample, and different runs share blind spots, so a miss in the spec can pass straight through all three. In <a href="the-future-of-programming-languages.html"><em>Musings on the Future of Programming Languages</em></a> I wondered whether reading code might become like reading the raw instructions a processor runs: possible, but rarely needed. If that happens, no person reads the code at all.</p>

<p>Computer scientists proposed a defence for this decades ago. In <a href="https://doi.org/10.1145/263699.263712">proof-carrying code</a>, described by George Necula in 1997, code from an untrusted source arrives with a mathematical proof that it follows an agreed set of safety rules, and the system receiving it checks the proof before running anything. Applied here, a model would produce the spec, the code and a proof that one matches the other, and a small, deliberately simple program called a proof checker would confirm the proof is valid. Trust would then rest on the proof checker, which is small enough for people to audit. It only protects what the spec describes, though. A tampered spec gets through, and so does an attack like the xz backdoor, which was hidden in the build process rather than in the code a proof would cover.</p>

<p>So code can get close to perfect and still need patching for as long as the world keeps changing. If machines take that patching over, the immune system becomes the thing to protect. Proof checking is one way to protect it, and it can only be as good as the spec.</p>]]></content><author><name>Cesaire Tobias</name></author><category term="ai" /><category term="software" /><category term="security" /><summary type="html"><![CDATA[If AI ends up writing near-perfect code, do software updates become new features only? No. Some change you choose and some the world forces on you, and AI may soon handle the forced kind on its own.]]></summary></entry><entry><title type="html">Lost in Translation: What Your Agent Gets When You Share a Document</title><link href="https://blog.sheetsolved.com/lost-in-translation.html" rel="alternate" type="text/html" title="Lost in Translation: What Your Agent Gets When You Share a Document" /><published>2026-09-17T00:00:00+00:00</published><updated>2026-09-17T00:00:00+00:00</updated><id>https://blog.sheetsolved.com/lost-in-translation</id><content type="html" xml:base="https://blog.sheetsolved.com/lost-in-translation.html"><![CDATA[<p>I used to write for people, but now more of the documents I share pass through a person on their way to an agent. I review an engineer’s work with Claude, the engineer reads the report (maybe), and then hands it to their own Claude session to make the changes. The engineer wants a page that reads well, and their session needs everything in the report intact. I’ve been writing in various flavours of Markdown for ages, and that’s been a really smooth workflow into this AI era. Sending a beautifully formatted HTML document has always been a no-brainer, but that was when the reader was exclusively human. I’m now starting to think about my non-human readers as well. So I tested a few things and came up with the following, which will all seem pretty intuitive to those of you who work with reproducible documentation flows.</p>

<h2 id="send-the-source">Send the source</h2>

<ul>
  <li>Send the Markdown, or whatever the page was rendered from. A person can read it as it is and render it if they want. It was the cheapest copy for Claude to read, and it’s the only one with no rendering step to lose anything in.</li>
  <li>Don’t send a link to a web page. In my tests the agent got only a short summary of the page, and once said it had read the whole thing.</li>
  <li>If Claude published the document as an artifact, ask it for the Markdown too and send that.</li>
  <li>Don’t send a single-file Quarto page. With only its Read tool, Claude couldn’t read it at all.</li>
  <li>Don’t rely on a PDF. Most of the tools that make one can break file paths across lines, and Claude doesn’t always put them back together. If you do print one from a browser, set inline code to <code class="language-plaintext highlighter-rouge">white-space: nowrap</code> first.</li>
</ul>

<h2 id="a-link">A link</h2>

<p>A pasted link to an ordinary web page goes through WebFetch, which by Claude Code’s own description converts the page to Markdown and has “a small fast model” answer a prompt about it. The agent gets that answer. I gave a session the arXiv link for <em>Attention Is All You Need</em> and asked it to read the whole page. It got back two short answers from a 189 KB page and replied that it had read it. In a second session on a newer Claude Code, it got the same kind of answers and said it couldn’t honestly claim to have read the page.</p>

<p>Claude can also publish a document as an artifact, a page on claude.ai you share by link. My review of my tidylearn package went out that way, with no Markdown copy. I tried the ways an agent could read it. Each copy was read in full, one session per copy, and the tokens each read added come from the session’s own usage records.</p>

<table>
  <thead>
    <tr>
      <th>Copy</th>
      <th>Tokens added</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Markdown copy, read from disk</td>
      <td>18.8K</td>
    </tr>
    <tr>
      <td>The page’s HTML, read from disk</td>
      <td>23.7K</td>
    </tr>
    <tr>
      <td>The link, read with the Artifact tool</td>
      <td>26.5K to 46.1K</td>
    </tr>
    <tr>
      <td>The link, read with WebFetch</td>
      <td>failed</td>
    </tr>
  </tbody>
</table>

<p>WebFetch didn’t get the page in any session I measured. Told to use it, it got HTTP 403 each time. Left to choose, the agent didn’t try, saying it couldn’t open claude.ai links, and asked for the content to be pasted in. Those were headless sessions (<code class="language-plaintext highlighter-rouge">claude -p</code>). In my own interactive session in VS Code, WebFetch opened the same link through my claude.ai login and returned it the way the Artifact tool does. So what WebFetch gets depends on the session, and I haven’t pinned down which part of it matters.</p>

<p>The tidylearn page is over 50 KB. The Artifact tool handed over the first 50 KB and saved the whole file to disk with an instruction to read all of it. In two sessions the agent then read the whole file, so most of the page went in twice. In a third, on a newer Claude Code, it read only the part it hadn’t seen. Two smaller artifacts of mine came back whole in one call. Having the Artifact tools available also adds context to every call a session makes, whether or not it reads an artifact.</p>

<p>I own this artifact, which is why the tool gave me its HTML. For an artifact someone else shared with you, the tool’s own description says a read returns an isolated summary. That’s the usual case for the engineer, and a summary can drop exactly the file paths and line numbers the engineer’s session needs.</p>

<p>Turning the page into Markdown after the fact was harder than it should be. The content isn’t in the HTML as text: a script builds it in the browser, so pandoc turned the saved page into an empty file. I had to render the page in a headless browser first. Asking Claude for the Markdown when it writes the document avoids all of that.</p>

<h2 id="rendered-copies">Rendered copies</h2>

<p>The other document is a review Claude wrote for a new product feature: 12 KB of Markdown with headings, lists, and file paths and line numbers in code formatting. I rendered it to HTML with pandoc, to a single-file HTML page with Quarto (<code class="language-plaintext highlighter-rouge">embed-resources</code>, which is what makes it one file you can send), and to PDF by printing the pandoc page from Chrome. Claude read each copy with its Read tool.</p>

<table>
  <thead>
    <tr>
      <th>Form</th>
      <th>File size</th>
      <th>Tokens added</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Markdown</td>
      <td>12 KB</td>
      <td>4.9K</td>
    </tr>
    <tr>
      <td>HTML, pandoc</td>
      <td>18 KB</td>
      <td>8.3K</td>
    </tr>
    <tr>
      <td>PDF, 5 pages</td>
      <td>161 KB</td>
      <td>12.5K</td>
    </tr>
    <tr>
      <td>HTML, Quarto single file</td>
      <td>1.2 MB</td>
      <td>not read</td>
    </tr>
  </tbody>
</table>

<p>With only its Read tool, Claude never got through the Quarto page. Quarto packs the page’s scripts, fonts and stylesheet into the file, some of them as single lines too long for the Read tool to return, and the report itself starts more than two thousand lines in. Claude hit one of those lines, couldn’t get past it, and said it could not read the file. Given search tools as well, it found where the report starts and read it from there, adding about three times the tokens the Markdown did.</p>

<p>The PDF broke file paths. Pandoc’s stylesheet lets inline code wrap, and Chrome treats a hyphen as a place to break a line, so a path like <code class="language-plaintext highlighter-rouge">load-model.ts:120-135</code> can end up split across two lines. Word and LibreOffice do the same. Typst breaks paths at slashes too, and LaTeX keeps them whole but lets a long one run off the edge of the page. Pull the text back out of a split path and you get a line break in the middle of it, or the hyphen gone: <code class="language-plaintext highlighter-rouge">loadmodel.ts</code>. The HTML kept every path exactly as written.</p>

<p>Claude usually put split paths back together, but not always. On two synthetic reviews with known paths, it gave every split path back exactly on one, and on the other some reads came back with a hyphen missing: <code class="language-plaintext highlighter-rouge">statement-run-8</code> as <code class="language-plaintext highlighter-rouge">statement-run8</code>. Adding <code class="language-plaintext highlighter-rouge">code { white-space: nowrap; }</code> to the page before printing from Chrome stopped the splits: each path moved whole to the next line.</p>

<h2 id="what-it-costs">What it costs</h2>

<p>For a light document the difference is small. The Markdown session above cost $0.070 at list price and the PDF session $0.101, against $0.046 for a session with no document. The PDF cost the most because a short PDF goes in as a document block, which <a href="https://platform.claude.com/docs/en/build-with-claude/pdf-support">Anthropic’s documentation</a> says carries each page as an image with its extracted text alongside, so Claude got it twice.</p>

<p>The dollar figures are Claude Code’s list prices, which is the unit it reports. On a subscription you won’t see a bill for any of this: the figures are a weighting for how much of your usage limit each copy takes up.</p>

<p>Whatever a read adds stays in the session, and every later call sends it again, from the prompt cache, until you <code class="language-plaintext highlighter-rouge">/clear</code>. Cache reads are billed at a fraction of normal input, so for a light document that stays small.</p>

<p>Heavy markup costs far more. The arXiv HTML page of <em>Attention Is All You Need</em> added 101.8K tokens against 19.4K for its PDF, and that session cost $0.59 at list price against $0.13.</p>

<h2 id="what-the-numbers-do-not-cover">What the numbers do not cover</h2>

<ul>
  <li>A link someone else shared with you. The artifact was read by the account that owns it, and the shared case, where the agent gets a summary, isn’t measured.</li>
  <li>Your own documents. A page with heavier markup than my reports will cost more to read, and a different renderer may break different things. The scripts in the footer run the same checks on your files.</li>
  <li>Other models. Everything ran on Sonnet 5 at low effort, and another model or effort level may repair broken paths more or less reliably.</li>
  <li>Other web pages. WebFetch was tried on one page, and another may come back differently.</li>
  <li>The measured sessions were headless. An interactive session can behave differently, as WebFetch did with the artifact link.</li>
  <li>The path checks show whether exact paths survive. They don’t show whether Claude understood the rest of a document.</li>
</ul>

<p><em>The rendered report and the paper were recorded on 16 September 2026 and everything else on 17 September, all on Claude Sonnet 5 at low effort, on Claude Code 2.1.272 except the link sessions repeated on 2.1.273. The WebFetch 403s, a second 2.1.272 Artifact-tool read of the tidylearn page (46.4K) and the interactive WebFetch call come from sessions that aren’t saved. The scripts and the saved runs are in the <a href="tokenwise-for-claude.html">tokenwise</a> repository, under <a href="https://github.com/ces0491/tokenwise/tree/main/experiments/doc-format">experiments/doc-format</a>, whose README has the figures left out here and the command to reprint each table for free. The report is a client document and isn’t published, and the artifacts are private, so their saved runs record hashes and versions but not the documents or links. <code class="language-plaintext highlighter-rouge">measure.mjs</code> measures your own files and links, <code class="language-plaintext highlighter-rouge">pdf-wrap.mjs</code> checks which paths survive each PDF tool without spending anything, and <code class="language-plaintext highlighter-rouge">recall.mjs</code> checks whether Claude gives the paths back. The 94 sessions behind the saved runs came to $9.96 at list price.</em></p>]]></content><author><name>Cesaire Tobias</name></author><category term="ai" /><category term="claude" /><category term="tooling" /><summary type="html"><![CDATA[Send an agent the source. In my tests a web link gave it a short summary, an artifact link failed or cost up to twice the page, a Quarto page defeated its Read tool, and PDFs broke file paths.]]></summary></entry><entry><title type="html">The Price of Asking: When a tokenwise Route Pays for Itself</title><link href="https://blog.sheetsolved.com/price-of-asking.html" rel="alternate" type="text/html" title="The Price of Asking: When a tokenwise Route Pays for Itself" /><published>2026-09-11T00:00:00+00:00</published><updated>2026-09-11T00:00:00+00:00</updated><id>https://blog.sheetsolved.com/price-of-asking</id><content type="html" xml:base="https://blog.sheetsolved.com/price-of-asking.html"><![CDATA[<p><a href="tokenwise-for-claude.html">tokenwise</a> answers one question for Claude Code: which model and effort level a piece of work needs. The answer comes from a model as well, so asking has a cost. On a small enough task that cost is more than the cheaper setting saves, and you would have done better picking a setting yourself.</p>

<p>I measured the cost of asking from three session settings and set it against what the cheaper settings saved on the benchmark behind the plugin’s routing table.</p>

<p><a href="https://blog.sheetsolved.com/assets/tokenwise/breakeven.svg"><img src="https://blog.sheetsolved.com/assets/tokenwise/breakeven.svg" alt="What each setting cost on four benchmark tasks, what a route costs from three settings, and the task sizes where a route pays for itself" /></a></p>

<h2 id="what-asking-costs">What asking costs</h2>

<p>A route runs on whatever model and effort your session is on, in a forked subagent, so only its answer comes back into the conversation.</p>

<table>
  <thead>
    <tr>
      <th>Asking from</th>
      <th>A route mid-session</th>
      <th>Context carried afterwards</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Sonnet 5, medium</td>
      <td>$0.04</td>
      <td>468 tokens</td>
    </tr>
    <tr>
      <td>Opus 5, high</td>
      <td>$0.10</td>
      <td>825 tokens</td>
    </tr>
    <tr>
      <td>Opus 5, xhigh</td>
      <td>$0.12</td>
      <td>822 tokens</td>
    </tr>
  </tbody>
</table>

<p>Claude Code’s model configuration docs put Max, Team Premium, Enterprise and API users on Opus 5 at high effort unless they change it, so the middle row is the default on those plans.</p>

<p>That context is carried on every later call until you clear. Anthropic’s pricing page lists an Opus 5 cache read at $0.50 per million tokens, so 825 tokens comes to $0.0004 a call.</p>

<p>Each session routed the same two short task descriptions. Sonnet 5 at medium recommended the same settings for both as the Opus sessions did, at less than half their cost. Its answers ran to 173 and 175 words against 308 to 371 for Opus.</p>

<h2 id="what-a-cheaper-setting-saves">What a cheaper setting saves</h2>

<p>Four of the benchmark’s tasks, all against a small invoicing library, ran on the expensive setting a routing claim was tested against and on the setting a route recommends. Every run in these cells passed its grader.</p>

<table>
  <thead>
    <tr>
      <th>Task</th>
      <th>Moved from</th>
      <th>Recommended</th>
      <th>Saved</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Implement a feature from a written spec</td>
      <td>Opus 5, xhigh: $1.65</td>
      <td>Sonnet 5, medium: $0.28</td>
      <td>$1.36, or 83%</td>
    </tr>
    <tr>
      <td>Rename a function across code, tests and README</td>
      <td>Opus 5, xhigh: $0.25</td>
      <td>Sonnet 5, low: $0.08</td>
      <td>$0.17, or 67%</td>
    </tr>
    <tr>
      <td>Fix a planted bug with a failing test</td>
      <td>Opus 5, high: $0.41</td>
      <td>Sonnet 5, medium: $0.12</td>
      <td>$0.29, or 70%</td>
    </tr>
    <tr>
      <td>Review a six-file diff</td>
      <td>Opus 5, high: $0.82</td>
      <td>Opus 5, low: $0.25</td>
      <td>$0.57, or 70%</td>
    </tr>
  </tbody>
</table>

<p>The median saving across the four is 70%.</p>

<h2 id="where-the-line-falls">Where the line falls</h2>

<p>A route is told to read only its own instructions and the one-line description you give it, so the size of the work should not change its cost, though the benchmark only priced it on short descriptions. What it saves grows with the work. With a 70% saving, a route breaks even on a task whose cost on your current setting is the route’s cost divided by 0.7, and it saves more than twice its cost only on tasks twice that size.</p>

<table>
  <thead>
    <tr>
      <th>Asking from</th>
      <th>Costs more than it saves below</th>
      <th>Saves less than twice its cost below</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Sonnet 5, medium</td>
      <td>$0.06</td>
      <td>$0.11</td>
    </tr>
    <tr>
      <td>Opus 5, high</td>
      <td>$0.14</td>
      <td>$0.29</td>
    </tr>
    <tr>
      <td>Opus 5, xhigh</td>
      <td>$0.17</td>
      <td>$0.33</td>
    </tr>
  </tbody>
</table>

<p>The bands are worked from the unrounded route costs, so dividing the rounded ones above can land a cent out.</p>

<p>The rename lands in the middle band. It cost $0.25 on Opus at xhigh, moving it to Sonnet at low saved $0.17, and asking cost $0.12, so asking came out $0.05 ahead. That is a gain too small to be worth the typing. The feature, the bug fix and the review each landed where a route saves more than twice its cost.</p>

<p>A route saves nothing when the session is already on the setting it recommends, and then its whole cost is lost whatever the size of the task.</p>

<p>In practice that comes down to three habits:</p>

<ul>
  <li>For a quick edit, pick a setting yourself: Sonnet at low effort, or Haiku.</li>
  <li>For a feature, a debugging session or a review, ask at the start of the phase, then <code class="language-plaintext highlighter-rouge">/clear</code> before you switch, so the switch does not re-process the conversation on the new setting. Switch with <code class="language-plaintext highlighter-rouge">s</code> in the <code class="language-plaintext highlighter-rouge">/model</code> picker: typing <code class="language-plaintext highlighter-rouge">/model &lt;name&gt;</code> also saves it as your default for new sessions.</li>
  <li>When you already know which setting the work needs, skip the question.</li>
</ul>

<h2 id="what-the-numbers-do-not-cover">What the numbers do not cover</h2>

<p>The benchmark’s tasks are small, and every saving above was measured at that size. The bands assume a larger task saves the same share on the cheaper setting, which is untested. If the cheaper setting fails and you run the task again, the retry comes off the saving.</p>

<p>The dollar figures are Claude Code’s list prices. On a subscription they are a weighting for comparing settings.</p>

<p>Changing model on a conversation already under way re-processes all of it on the new model, and so does changing effort on most models, according to Claude Code’s prompt caching docs. Straight after a <code class="language-plaintext highlighter-rouge">/clear</code>, only the system prompt and project context are re-processed, as in a new session. A switch without one adds a cost these figures leave out.</p>

<p>The route costs come from one run per setting. The 1.1.1 route, run twice on identical text, cost $0.155 and $0.129, so two runs of one route can differ by $0.025, against the rename’s $0.05 margin. Sonnet 5 at high, the default on Pro and Team Standard plans, was not measured.</p>

<p><em>Task runs on Claude Code 2.1.263, 8 September 2026. Routes on tokenwise 1.3.0 and Claude Code 2.1.272. The code, the transcripts and the chart’s generator are at <a href="https://github.com/ces0491/tokenwise">github.com/ces0491/tokenwise</a>, and two commands check the figures without spending anything: <code class="language-plaintext highlighter-rouge">node bench/breakeven.mjs --check</code> confirms the chart follows from the committed runs, and <code class="language-plaintext highlighter-rouge">node bench/skill-cost.mjs --compare 1.3.0,1.3.0-opus-high@1.3.0,1.3.0-sonnet-medium@1.3.0</code> prints the route costs from the saved sessions.</em></p>]]></content><author><name>Cesaire Tobias</name></author><category term="ai" /><category term="claude" /><category term="tooling" /><summary type="html"><![CDATA[A tokenwise route has a cost of its own. What asking costs from three Claude Code settings, and the task size below which the answer costs more than it saves.]]></summary></entry><entry><title type="html">tokenwise for Claude Code</title><link href="https://blog.sheetsolved.com/tokenwise-for-claude.html" rel="alternate" type="text/html" title="tokenwise for Claude Code" /><published>2026-09-09T00:00:00+00:00</published><updated>2026-09-09T00:00:00+00:00</updated><id>https://blog.sheetsolved.com/tokenwise-for-claude</id><content type="html" xml:base="https://blog.sheetsolved.com/tokenwise-for-claude.html"><![CDATA[<p>On a subscription the money is settled before the month starts. What you spend from then on is your usage limit, and it goes on whatever you happen to point Claude Code at, including phases that would have finished on a cheaper setting.</p>

<p>Every API call re-sends the whole conversation. What a task consumes is context size multiplied by the number of calls, and the model and effort level you pick set the rate on top of that.</p>

<p>On my own Claude Code transcripts the input side runs 444 tokens for every one of output, and almost all of it is conversation being re-read from cache. <code class="language-plaintext highlighter-rouge">node bench/context-profile.mjs</code> in the repository below produces that table from your transcripts instead of mine.</p>

<p>That leaves two levers. Context size is mostly a discipline problem: when you clear, what you delegate to a subagent, how large a slice you take on at once. tokenwise is for the other one, the model and effort you pick — a decision made once at the start and then usually left alone for the rest of the session.</p>

<h2 id="what-it-does">What it does</h2>

<p>Its core is one skill:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>/tokenwise:route implement the plan in docs/plan.md, about 12 files
</code></pre></div></div>

<p>The answer names a model and an effort level and how to switch to them for the session, says whether to clear the context first and what the switch costs if you don’t, says what to push into subagents, says what the cheaper choice gives up, and gives you one check to run to know the phase is finished before you trust it.</p>

<p>Every row of the routing table underneath it names where to start, and the seven rows with somewhere to escalate to name what has to go wrong before you move. The escalation rule comes from Anthropic’s own guidance: if the model failed with the context it had, it did not know enough, so change the model; if it skipped files or stopped early, it did not try hard enough, so raise the effort. I used to treat those as one problem.</p>

<p>Reviewing a diff starts on the expensive model at low effort. High effort on a diff you can hold in your head buys more turns spent re-reading it, and the bound on what a review can usefully do is the subject of <a href="bounding-ai-code-reviews.html">How Long Is a Piece of String?</a>.</p>

<p>A second skill, <code class="language-plaintext highlighter-rouge">/tokenwise:setup</code>, makes Sonnet 5 at medium the default for new sessions if you say yes. Two hooks warn when resuming a conversation or switching model is about to re-send it to an empty cache, and write nothing into the conversation.</p>

<h2 id="what-it-will-not-do">What it will not do</h2>

<p>It cannot switch anything for you. Nothing can change a running session’s model or effort level — that is <code class="language-plaintext highlighter-rouge">/model</code> and <code class="language-plaintext highlighter-rouge">/effort</code>, typed by you. The skill recommends, and tells you what the recommendation gives up.</p>

<p>What the plugin does add to every session is the route skill’s description, 153 tokens, so that Claude knows the skill is there. That also means Claude can run the skill without being told to, and in testing it did when a session asked in plain words which model and effort to use.</p>

<h2 id="what-it-does-not-know">What it does not know</h2>

<p>Five of the nine rows in the routing table carry no measurement behind them, and each says so in the table. They are informed guesses about work I have not benchmarked.</p>

<p>The four that are measured come from a benchmark I built to check the table, and it changed the table. Four of the eight claims I had written in turned out to be wrong: twice by recommending a setting more expensive than the work needed, once by overstating a saving, and once on a piece of process that charges you for writing a plan down and reading it back. That is its own story and I have written it up separately.</p>

<p>The benchmark counts in dollars, because list prices are the unit Claude Code reports. On a subscription they measure how much of your month a task eats rather than a charge you will see.</p>

<p>Every graded run passed, whatever the setting, so the benchmark compares consumption at equal outcomes and never reaches the point where an expensive setting earns its keep by succeeding where a cheap one fails. The graders, the pre-registered pass criteria and the raw results for every run are in the repository, so you can disagree with me using my own data.</p>

<h2 id="install">Install</h2>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>/plugin marketplace add ces0491/tokenwise
/plugin install tokenwise@ces0491-plugins
</code></pre></div></div>

<p>Or read it before you install it:</p>

<div class="language-sh highlighter-rouge"><div class="highlight"><pre class="highlight"><code>git clone https://github.com/ces0491/tokenwise <span class="o">&amp;&amp;</span> <span class="nb">cd </span>tokenwise
node bench/summarize.mjs <span class="nt">--check</span>    <span class="c"># do the published tables follow from the published runs?</span>
node bench/context-profile.mjs      <span class="c"># the token figures above, against your own transcripts</span>
</code></pre></div></div>

<p>Neither makes an API call, so your limit is untouched. MIT licensed, at <a href="https://github.com/ces0491/tokenwise">github.com/ces0491/tokenwise</a>.</p>]]></content><author><name>Cesaire Tobias</name></author><category term="ai" /><category term="claude" /><category term="tooling" /><summary type="html"><![CDATA[On a subscription the cost of a task is a share of your usage limit. A Claude Code plugin for not spending that share on work that never needed it.]]></summary></entry><entry><title type="html">Pulling on Threads: Scoring Bands I Never Checked</title><link href="https://blog.sheetsolved.com/scoring-bands-i-never-checked.html" rel="alternate" type="text/html" title="Pulling on Threads: Scoring Bands I Never Checked" /><published>2026-09-04T00:00:00+00:00</published><updated>2026-09-04T00:00:00+00:00</updated><id>https://blog.sheetsolved.com/scoring-bands-i-never-checked</id><content type="html" xml:base="https://blog.sheetsolved.com/scoring-bands-i-never-checked.html"><![CDATA[<p>I have a dashboard that <a href="is-this-package-safe-to-depend-on.html">scores CRAN packages</a> on how reasonable they may be to include in your next project. Monthly downloads are worth 20 of its 100 points, and I banded them by powers of ten: under 100, then 1,000, then 10,000, then 100,000. Five bands across four orders of magnitude, for a quantity that runs from a handful to millions. It looked about right, and I didn’t give it any further thought.</p>

<p>What sent me back to it was a different bug. The momentum factor had been dividing a package’s lifetime downloads by twelve months whether or not it had been on CRAN that long, so a package four months old got a baseline three times too low and its decline came out as growth. Pulling on that thread began to expose a few other gaps.</p>

<p>So: 600 packages spaced evenly through the 24,887 CRAN publishes, in name order, and cranlogs for last month’s totals in three batched requests.</p>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">idx</span><span class="w">  </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">available.packages</span><span class="p">(</span><span class="n">repos</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"https://cloud.r-project.org"</span><span class="p">,</span><span class="w">
                           </span><span class="n">filters</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">character</span><span class="p">())</span><span class="w">
</span><span class="n">pkgs</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">sort</span><span class="p">(</span><span class="n">rownames</span><span class="p">(</span><span class="n">idx</span><span class="p">))</span><span class="w">
</span><span class="n">sel</span><span class="w">  </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">unique</span><span class="p">(</span><span class="n">pkgs</span><span class="p">[</span><span class="nf">round</span><span class="p">(</span><span class="n">seq</span><span class="p">(</span><span class="m">1</span><span class="p">,</span><span class="w"> </span><span class="nf">length</span><span class="p">(</span><span class="n">pkgs</span><span class="p">),</span><span class="w"> </span><span class="n">length.out</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="m">600</span><span class="p">))])</span><span class="w">

</span><span class="n">counts</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">unlist</span><span class="p">(</span><span class="n">lapply</span><span class="p">(</span><span class="n">split</span><span class="p">(</span><span class="n">sel</span><span class="p">,</span><span class="w"> </span><span class="nf">ceiling</span><span class="p">(</span><span class="nf">seq_along</span><span class="p">(</span><span class="n">sel</span><span class="p">)</span><span class="w"> </span><span class="o">/</span><span class="w"> </span><span class="m">200</span><span class="p">)),</span><span class="w"> </span><span class="k">function</span><span class="p">(</span><span class="n">ch</span><span class="p">)</span><span class="w"> </span><span class="p">{</span><span class="w">
  </span><span class="n">jsonlite</span><span class="o">::</span><span class="n">fromJSON</span><span class="p">(</span><span class="n">paste0</span><span class="p">(</span><span class="w">
    </span><span class="s2">"https://cranlogs.r-pkg.org/downloads/total/last-month/"</span><span class="p">,</span><span class="w">
    </span><span class="n">paste</span><span class="p">(</span><span class="n">ch</span><span class="p">,</span><span class="w"> </span><span class="n">collapse</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">","</span><span class="p">)))</span><span class="o">$</span><span class="n">downloads</span><span class="w">
</span><span class="p">}),</span><span class="w"> </span><span class="n">use.names</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="kc">FALSE</span><span class="p">)</span><span class="w">

</span><span class="n">median</span><span class="p">(</span><span class="n">counts</span><span class="p">)</span><span class="w">
</span><span class="c1">#&gt; [1] 277.5</span><span class="w">
</span><span class="nf">sum</span><span class="p">(</span><span class="n">counts</span><span class="w"> </span><span class="o">&gt;=</span><span class="w"> </span><span class="m">100</span><span class="w"> </span><span class="o">&amp;</span><span class="w"> </span><span class="n">counts</span><span class="w"> </span><span class="o">&lt;</span><span class="w"> </span><span class="m">1000</span><span class="p">)</span><span class="w">
</span><span class="c1">#&gt; [1] 532</span><span class="w">
</span><span class="nf">min</span><span class="p">(</span><span class="n">counts</span><span class="p">)</span><span class="w">
</span><span class="c1">#&gt; [1] 100</span><span class="w">
</span></code></pre></div></div>

<p>532 of the 600 in one band. The median package gets 277 downloads a month. The factor was carrying a fifth of the weight and telling me almost nothing: a package on 150 a month sits at the 4th percentile of that sample, one on 900 sits at the 88th, and both scored 6 out of 20.</p>

<p>The part I hadn’t gone looking for was the band underneath. Labelled <em>very low</em>, and empty. The smallest count anywhere in 600 packages was exactly 100.</p>

<p>It turns out that cranlogs counts fetches from Posit’s mirror, and crawlers and mirror syncs get counted alongside installs. The raw logs record the R version a client reported, and a fetch with none didn’t come from <code class="language-plaintext highlighter-rouge">install.packages()</code>. I sampled 9 packages over 13 days: the share reporting a version ran from 9% to 54%, lowest for the packages with the fewest downloads. Nothing on CRAN reads as zero, and the smaller the package, the less of its count is people.</p>

<p>Then I went looking for prior art, in the wrong order. Peter Li shipped <a href="https://cran.r-project.org/package=packageRank">packageRank</a> in 2019 and <a href="https://www.r-bloggers.com/2020/05/counting-and-visualizing-cran-downloads-with-packagerank-with-caveats/">wrote the caveat up</a> in 2020, working from the raw logs rather than the aggregates. Had I read that before writing the scoring code, I’d have started from his caveats instead of reconstructing a rougher version of them.</p>

<p>The bands now break at 50, 200, 500, 2,000, 10,000 and 100,000, cut finer through the range the packages are actually in and left as absolute counts: a few hundred downloads a month is a small user base however much of the repository it beats.</p>

<p>The same pass turned up two more. Maturity was using release count as a proxy for age, and the two correlate at 0.49. Having no reverse dependencies was rendering as a red flag, when 17,313 of the 24,887 packages on CRAN have none.</p>

<p>All of it had been live since the end of March, getting the obvious cases right. ggplot2 near the top, an archived package near the bottom, which seemed like common-sense sniff tests. The errors sat in the middle of the distribution, which is the only part anyone really needs a score for. Sometimes a sniff test is good enough — not everything you build needs to be bulletproof — but I’ve been shown time and time again to never stop pulling on the threads.</p>

<p><em>Figures from a live run on 4 September 2026; cranlogs totals move daily.</em></p>]]></content><author><name>Cesaire Tobias</name></author><category term="R" /><category term="r" /><category term="cran" /><category term="scoring" /><summary type="html"><![CDATA[I banded monthly downloads by powers of ten without checking where the packages actually were. A sample of 600 put 532 of them in one band, and the band below caught nothing at all.]]></summary></entry><entry><title type="html">Is This Package Safe to Depend On?</title><link href="https://blog.sheetsolved.com/is-this-package-safe-to-depend-on.html" rel="alternate" type="text/html" title="Is This Package Safe to Depend On?" /><published>2026-08-22T00:00:00+00:00</published><updated>2026-08-22T00:00:00+00:00</updated><id>https://blog.sheetsolved.com/is-this-package-safe-to-depend-on</id><content type="html" xml:base="https://blog.sheetsolved.com/is-this-package-safe-to-depend-on.html"><![CDATA[<p>Adding a package to <code class="language-plaintext highlighter-rouge">DESCRIPTION</code> takes a few seconds. Taking one back out, three years later, can take weeks (probably just days with some AI assistance).</p>

<p>The decision to depend on something is made in a moment, usually while you’re focused on something else entirely — you need a date parser, someone on Stack Overflow (remember Stack Overflow?) used this one, it works, move on. But every so often, the package stops building against the current R release, or the maintainer’s email starts bouncing, or CRAN archives it and that glossed-over reverse dependency becomes a proper headache.</p>

<p>We’ve built good habits around a lot of things that are cheaper to get wrong than this. We review pull requests that change five lines but we don’t review the line that adopts twelve thousand lines of somebody else’s code. Which probably says more about human nature than the explicit habits of coders, but I digress.</p>

<h2 id="cran-already-publishes-the-answer">CRAN already publishes the answer</h2>

<p>The information you’d want is public, free, and machine-readable. It’s just spread across three separate services, none of which is the page you land on when you Google a package name.</p>

<ul>
  <li><a href="https://crandb.r-pkg.org"><strong>crandb</strong></a> has package metadata, the full release timeline, reverse dependency listings (with a caveat I’ll come back to), and — importantly — whether the package has ever been archived.</li>
  <li><a href="https://cranlogs.r-pkg.org"><strong>cranlogs</strong></a> has download counts, daily or aggregated over arbitrary windows.</li>
  <li><a href="https://search.r-pkg.org"><strong>search.r-pkg.org</strong></a> does full-text search across all of it.</li>
</ul>

<p>Most of us don’t check all three before adding a dependency. At least, I didn’t. I was fortunate enough to have good training that made me think about these sorts of things but the overhead of a thorough package audit was just too boring and honestly a bit hand-wavy — if there was a good enough package to do a job you needed, you used it. But we have AI for these boring jobs now (like writing tests — I’ve never seen this many tests in repos ever [which is great btw]) so there’s no longer an excuse to implement the good practices you know you should be but don’t.</p>

<h2 id="the-five-signals">The five signals</h2>

<p>Five signals are worth weighing. Each is easy to over-read on its own, which matters more than the list.</p>

<p><strong>Recency.</strong> When was the last CRAN release? This is the signal most reach for but is also easy to misinterpret. A package with no release in three years might be abandoned — or it might be <em>finished</em>. Some of the best small packages on CRAN do one thing, do it correctly, and have had no reason to change. Recency only becomes damning in combination: a long gap <em>plus</em> failing checks on current R is a package on its way to the archive. A long gap plus clean checks is often just a stable package.</p>

<p><strong>Download momentum.</strong> What matters is the direction. A package whose downloads have halved over a year is telling you something that its lifetime total is hiding. It is the most forward-looking of the five, and the easiest to compute wrong. cranlogs returns rows only from a package’s first publication, so a four-month-old package has four months of data. Divide its total by twelve to get a baseline and you understate that baseline threefold, which turns a decline into growth. Compare mean downloads per day over the last 30 days against the mean per day over the days before them, and report nothing at all with fewer than 60 days of history.</p>

<p><strong>Download volume.</strong> Useful, but inflated. cranlogs counts fetches from Posit’s mirror, so it is a sample of CRAN, and crawlers and mirror syncs are counted alongside installs. Installing a package also pulls its dependencies down with it, which means anything sitting in a popular dependency tree posts numbers that reflect <em>its dependents’</em> popularity and not necessarily its own adoption. Peter Li made that point, and the observation that the bias dilutes as real downloads rise, in <a href="https://cran.r-project.org/package=packageRank">packageRank</a> and <a href="https://www.r-bloggers.com/2020/05/counting-and-visualizing-cran-downloads-with-packagerank-with-caveats/">his write-up of it</a>. A package with two million downloads a month might have very few humans who chose it deliberately.</p>

<p><strong>Ecosystem adoption.</strong> How many other CRAN packages depend on this one. It functions as a form of insurance: CRAN is reluctant to archive a package that would break a long tail of dependents, and maintainers of widely-depended-on packages get pressure, help, and sometimes successors in a way that solo packages don’t. High reverse dependency counts mean the ecosystem itself has a stake in the thing continuing to work.</p>

<p><strong>Maturity.</strong> How long the package has been on CRAN. This separates something first published in 2007 from something first published six months ago, which the other four signals can’t distinguish. It’s the weakest of the five, and weighted accordingly, but it captures the case where seemingly promising new packages are relegated on the author’s list of priorities.</p>

<p>Using release count as proxy for age proves a poor decision. Across a systematic sample of 140 CRAN packages the two correlate at 0.49, and 17 of them had been on CRAN eight years or more with four releases or fewer — stable packages that a release-count measure files as new. Take the age from the first publication date, and show the release count beside it as a track record.</p>

<h2 id="doing-it-yourself">Doing it yourself</h2>

<p>You don’t need an app for this. Two API calls and one base R function get you most of the way:</p>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">library</span><span class="p">(</span><span class="n">httr2</span><span class="p">)</span><span class="w">

</span><span class="c1"># The slow part; available.packages() caches for the session.</span><span class="w">
</span><span class="n">cran_db</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">available.packages</span><span class="p">(</span><span class="n">repos</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"https://cloud.r-project.org"</span><span class="p">)</span><span class="w">

</span><span class="n">cran_signals</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="k">function</span><span class="p">(</span><span class="n">pkg</span><span class="p">,</span><span class="w"> </span><span class="n">db</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">cran_db</span><span class="p">)</span><span class="w"> </span><span class="p">{</span><span class="w">
  </span><span class="n">meta</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">request</span><span class="p">(</span><span class="s2">"https://crandb.r-pkg.org"</span><span class="p">)</span><span class="w"> </span><span class="o">|&gt;</span><span class="w">
    </span><span class="n">req_url_path_append</span><span class="p">(</span><span class="n">pkg</span><span class="p">,</span><span class="w"> </span><span class="s2">"all"</span><span class="p">)</span><span class="w"> </span><span class="o">|&gt;</span><span class="w">
    </span><span class="n">req_perform</span><span class="p">()</span><span class="w"> </span><span class="o">|&gt;</span><span class="w">
    </span><span class="n">resp_body_json</span><span class="p">()</span><span class="w">

  </span><span class="n">dl</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">request</span><span class="p">(</span><span class="s2">"https://cranlogs.r-pkg.org"</span><span class="p">)</span><span class="w"> </span><span class="o">|&gt;</span><span class="w">
    </span><span class="n">req_url_path_append</span><span class="p">(</span><span class="s2">"downloads"</span><span class="p">,</span><span class="w"> </span><span class="s2">"total"</span><span class="p">,</span><span class="w"> </span><span class="s2">"last-month"</span><span class="p">,</span><span class="w"> </span><span class="n">pkg</span><span class="p">)</span><span class="w"> </span><span class="o">|&gt;</span><span class="w">
    </span><span class="n">req_perform</span><span class="p">()</span><span class="w"> </span><span class="o">|&gt;</span><span class="w">
    </span><span class="n">resp_body_json</span><span class="p">()</span><span class="w">

  </span><span class="n">revdeps</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">tools</span><span class="o">::</span><span class="n">package_dependencies</span><span class="p">(</span><span class="w">
    </span><span class="n">pkg</span><span class="p">,</span><span class="w"> </span><span class="n">db</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">db</span><span class="p">,</span><span class="w"> </span><span class="n">reverse</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="kc">TRUE</span><span class="p">,</span><span class="w">
    </span><span class="n">which</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="nf">c</span><span class="p">(</span><span class="s2">"Depends"</span><span class="p">,</span><span class="w"> </span><span class="s2">"Imports"</span><span class="p">,</span><span class="w"> </span><span class="s2">"LinkingTo"</span><span class="p">)</span><span class="w">
  </span><span class="p">)[[</span><span class="m">1</span><span class="p">]]</span><span class="w">

  </span><span class="n">dates</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">sort</span><span class="p">(</span><span class="n">unlist</span><span class="p">(</span><span class="n">meta</span><span class="o">$</span><span class="n">timeline</span><span class="p">))</span><span class="w">

  </span><span class="nf">list</span><span class="p">(</span><span class="w">
    </span><span class="n">version</span><span class="w">       </span><span class="o">=</span><span class="w"> </span><span class="n">meta</span><span class="o">$</span><span class="n">latest</span><span class="p">,</span><span class="w">
    </span><span class="n">first_release</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">unname</span><span class="p">(</span><span class="n">as.Date</span><span class="p">(</span><span class="n">dates</span><span class="p">[</span><span class="m">1</span><span class="p">])),</span><span class="w">
    </span><span class="n">last_release</span><span class="w">  </span><span class="o">=</span><span class="w"> </span><span class="n">unname</span><span class="p">(</span><span class="n">as.Date</span><span class="p">(</span><span class="n">dates</span><span class="p">[</span><span class="nf">length</span><span class="p">(</span><span class="n">dates</span><span class="p">)])),</span><span class="w">
    </span><span class="n">releases</span><span class="w">      </span><span class="o">=</span><span class="w"> </span><span class="nf">length</span><span class="p">(</span><span class="n">dates</span><span class="p">),</span><span class="w">
    </span><span class="n">revdeps</span><span class="w">       </span><span class="o">=</span><span class="w"> </span><span class="nf">length</span><span class="p">(</span><span class="n">revdeps</span><span class="p">),</span><span class="w">
    </span><span class="n">archived</span><span class="w">      </span><span class="o">=</span><span class="w"> </span><span class="n">meta</span><span class="o">$</span><span class="n">archived</span><span class="p">,</span><span class="w">
    </span><span class="n">downloads_30d</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">dl</span><span class="p">[[</span><span class="m">1</span><span class="p">]]</span><span class="o">$</span><span class="n">downloads</span><span class="w">
  </span><span class="p">)</span><span class="w">
</span><span class="p">}</span><span class="w">

</span><span class="n">str</span><span class="p">(</span><span class="n">cran_signals</span><span class="p">(</span><span class="s2">"ggplot2"</span><span class="p">))</span><span class="w">
</span></code></pre></div></div>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>List of 7
 $ version      : chr "4.0.3"
 $ first_release: Date[1:1], format: "2007-06-01"
 $ last_release : Date[1:1], format: "2026-04-22"
 $ releases     : int 55
 $ revdeps      : int 4833
 $ archived     : logi FALSE
 $ downloads_30d: int 2000824
</code></pre></div></div>

<p>Nineteen years, 55 releases, two million downloads a month, 4,833 packages depending on it, never archived. You didn’t need the data to know ggplot2 is safe — if your method disagrees with the obvious cases, you probably need to go back and check your working.</p>

<p>The <code class="language-plaintext highlighter-rouge">archived</code> field is worth singling out. A package that has been archived and restored has already demonstrated the failure mode you’re worried about, and it’s the one signal here that’s binary rather than a matter of degree.</p>

<h2 id="crandb-has-two-reverse-dependency-numbers-and-one-of-them-is-a-trap">crandb has two reverse dependency numbers and one of them is a trap</h2>

<p>The reverse dependency count above doesn’t come from the same call as everything else - deliberately.</p>

<p>The <code class="language-plaintext highlighter-rouge">/{package}/all</code> response carries a <code class="language-plaintext highlighter-rouge">revdeps</code> field. It arrives in the JSON you already have, it’s a single integer, and it is very tempting. It also isn’t the number you want:</p>

<table>
  <thead>
    <tr>
      <th>Package</th>
      <th><code class="language-plaintext highlighter-rouge">revdeps</code> field</th>
      <th>Actual strong revdeps</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>ggplot2</td>
      <td>410</td>
      <td>4,833</td>
    </tr>
    <tr>
      <td>dplyr</td>
      <td>63</td>
      <td>4,916</td>
    </tr>
    <tr>
      <td>jsonlite</td>
      <td>63</td>
      <td>1,653</td>
    </tr>
    <tr>
      <td>httr2</td>
      <td><em>absent</em></td>
      <td>455</td>
    </tr>
  </tbody>
</table>

<p>dplyr and jsonlite both come back with 63, though dplyr has three times as many actual dependents, and httr2 has no value at all. Whatever that field counts, it isn’t what its name suggests, and a score weighted on it would be ranking packages by an artefact.</p>

<p>The number you actually want is at a different crandb endpoint, <code class="language-plaintext highlighter-rouge">/-/revdeps/{package}</code>, which returns the dependent package names split by relationship:</p>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">revdeps</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">request</span><span class="p">(</span><span class="s2">"https://crandb.r-pkg.org"</span><span class="p">)</span><span class="w"> </span><span class="o">|&gt;</span><span class="w">
  </span><span class="n">req_url_path_append</span><span class="p">(</span><span class="s2">"-"</span><span class="p">,</span><span class="w"> </span><span class="s2">"revdeps"</span><span class="p">,</span><span class="w"> </span><span class="s2">"ggplot2"</span><span class="p">)</span><span class="w"> </span><span class="o">|&gt;</span><span class="w">
  </span><span class="n">req_perform</span><span class="p">()</span><span class="w"> </span><span class="o">|&gt;</span><span class="w">
  </span><span class="n">resp_body_json</span><span class="p">()</span><span class="w">

</span><span class="n">lengths</span><span class="p">(</span><span class="n">revdeps</span><span class="o">$</span><span class="n">ggplot2</span><span class="p">)</span><span class="w">
</span><span class="c1">#&gt;   Imports  Suggests   Depends  Enhances</span><span class="w">
</span><span class="c1">#&gt;      4440      2033       395         7</span><span class="w">
</span></code></pre></div></div>

<p>That agrees with <code class="language-plaintext highlighter-rouge">tools::package_dependencies()</code> to within a handful of packages — 4,835 against 4,833 — and the difference is just which mirror snapshot each one saw. Either route is correct. The single integer in the metadata blob is not.</p>

<p>The discrepancy is only visible if you already have a sense of the right magnitude. When you assemble a metric from convenient sources, the field that’s easiest to reach is not always the field you want, and a plausible wrong number is much harder to notice than a missing one.</p>

<h2 id="the-score">The Score</h2>

<p>Five signals is four more than most people will weigh in the few seconds the decision actually gets. So I collapsed them into a 0–100 viability score, weighted like this:</p>

<table>
  <thead>
    <tr>
      <th>Factor</th>
      <th>Weight</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Recency</td>
      <td>30%</td>
    </tr>
    <tr>
      <td>Download momentum</td>
      <td>25%</td>
    </tr>
    <tr>
      <td>Download volume</td>
      <td>20%</td>
    </tr>
    <tr>
      <td>Ecosystem adoption</td>
      <td>15%</td>
    </tr>
    <tr>
      <td>Maturity</td>
      <td>10%</td>
    </tr>
  </tbody>
</table>

<p>Those weights are a judgement call. Recency and momentum are the forward-looking signals and volume is the backward-looking one, so the front of the list gets the weight — but somebody with different priorities should weight them differently.</p>

<p>The score is a triage tool. It’s good for sorting a shortlist and for noticing that something you assumed was fine has been declining. It is not a verdict, and a low score on a small, finished, well-written package that does exactly what you need is probably a false alarm you should overrule.</p>

<h2 id="the-app">The app</h2>

<p>I built <a href="https://github.com/ces0491/cranExploreR"><strong>cranExploreR</strong></a>. It’s a Shiny app — <code class="language-plaintext highlighter-rouge">bslib</code>, <code class="language-plaintext highlighter-rouge">plotly</code>, <code class="language-plaintext highlighter-rouge">httr2</code> — that pulls the signals above, draws the twelve-month download trend, and lets you put two or three candidate packages side by side, which is usually the actual question.</p>

<p>It’s <a href="https://019d3e9e-b1a7-77dc-9266-40ce0b717eb3.share.connect.posit.cloud/">running here</a>. Source code on GitHub.</p>

<p>The app is convenient but the question is answerable from data CRAN has been publishing all along.</p>

<p>The tooling is there too. <a href="https://cran.r-project.org/package=packageRank">packageRank</a> has been on CRAN since 2019, across 25 releases, and treats the download half of this more carefully than anything above — it works from the raw logs rather than the aggregates, which is what it takes to separate an install from a crawler. The data and the tools were both available the whole time. But the percieved overhead of maintaining the discipline of checking meant I wasn’t always doing this. I hope that cranExploreR makes the check trivial enough for it to always be a part of any of your consequential workflows.</p>]]></content><author><name>Cesaire Tobias</name></author><category term="R" /><category term="r" /><category term="cran" /><category term="dependencies" /><category term="shiny" /><summary type="html"><![CDATA[Adding a dependency takes seconds and removing one can take weeks. CRAN already publishes enough to answer the question before you commit — it's just spread across three APIs most of us don't check.]]></summary></entry><entry><title type="html">From Model to Report: How tidylearn Simplifies ML Reporting</title><link href="https://blog.sheetsolved.com/tidylearn-reporting.html" rel="alternate" type="text/html" title="From Model to Report: How tidylearn Simplifies ML Reporting" /><published>2026-08-21T00:00:00+00:00</published><updated>2026-08-21T00:00:00+00:00</updated><id>https://blog.sheetsolved.com/tidylearn-reporting</id><content type="html" xml:base="https://blog.sheetsolved.com/tidylearn-reporting.html"><![CDATA[<style>.gt-post-table table {
  font-family: system-ui, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif, 'Apple Color Emoji', 'Segoe UI Emoji', 'Segoe UI Symbol', 'Noto Color Emoji';
  -webkit-font-smoothing: antialiased;
  -moz-osx-font-smoothing: grayscale;
}

.gt-post-table thead, .gt-post-table tbody, .gt-post-table tfoot, .gt-post-table tr, .gt-post-table td, .gt-post-table th {
  border-style: none;
}

.gt-post-table p {
  margin: 0;
  padding: 0;
}

.gt-post-table .gt_table {
  display: table;
  border-collapse: collapse;
  line-height: normal;
  margin-left: auto;
  margin-right: auto;
  color: #333333;
  font-size: 13px;
  font-weight: normal;
  font-style: normal;
  background-color: #FFFFFF;
  width: auto;
  border-top-style: solid;
  border-top-width: 2px;
  border-top-color: #2C3E50;
  border-right-style: none;
  border-right-width: 2px;
  border-right-color: #D3D3D3;
  border-bottom-style: solid;
  border-bottom-width: 2px;
  border-bottom-color: #2C3E50;
  border-left-style: none;
  border-left-width: 2px;
  border-left-color: #D3D3D3;
}

.gt-post-table .gt_caption {
  padding-top: 4px;
  padding-bottom: 4px;
}

.gt-post-table .gt_title {
  color: #FFFFFF;
  font-size: 16px;
  font-weight: initial;
  padding-top: 4px;
  padding-bottom: 4px;
  padding-left: 5px;
  padding-right: 5px;
  border-bottom-color: #FFFFFF;
  border-bottom-width: 0;
}

.gt-post-table .gt_subtitle {
  color: #FFFFFF;
  font-size: 12px;
  font-weight: initial;
  padding-top: 3px;
  padding-bottom: 5px;
  padding-left: 5px;
  padding-right: 5px;
  border-top-color: #FFFFFF;
  border-top-width: 0;
}

.gt-post-table .gt_heading {
  background-color: #2C3E50;
  text-align: center;
  border-bottom-color: #FFFFFF;
  border-left-style: none;
  border-left-width: 1px;
  border-left-color: #D3D3D3;
  border-right-style: none;
  border-right-width: 1px;
  border-right-color: #D3D3D3;
}

.gt-post-table .gt_bottom_border {
  border-bottom-style: solid;
  border-bottom-width: 2px;
  border-bottom-color: #D3D3D3;
}

.gt-post-table .gt_col_headings {
  border-top-style: solid;
  border-top-width: 2px;
  border-top-color: #D3D3D3;
  border-bottom-style: solid;
  border-bottom-width: 2px;
  border-bottom-color: #D3D3D3;
  border-left-style: none;
  border-left-width: 1px;
  border-left-color: #D3D3D3;
  border-right-style: none;
  border-right-width: 1px;
  border-right-color: #D3D3D3;
}

.gt-post-table .gt_col_heading {
  color: #FFFFFF;
  background-color: #34495E;
  font-size: 100%;
  font-weight: bold;
  text-transform: inherit;
  border-left-style: none;
  border-left-width: 1px;
  border-left-color: #D3D3D3;
  border-right-style: none;
  border-right-width: 1px;
  border-right-color: #D3D3D3;
  vertical-align: bottom;
  padding-top: 5px;
  padding-bottom: 6px;
  padding-left: 5px;
  padding-right: 5px;
  overflow-x: hidden;
}

.gt-post-table .gt_column_spanner_outer {
  color: #FFFFFF;
  background-color: #34495E;
  font-size: 100%;
  font-weight: bold;
  text-transform: inherit;
  padding-top: 0;
  padding-bottom: 0;
  padding-left: 4px;
  padding-right: 4px;
}

.gt-post-table .gt_column_spanner_outer:first-child {
  padding-left: 0;
}

.gt-post-table .gt_column_spanner_outer:last-child {
  padding-right: 0;
}

.gt-post-table .gt_column_spanner {
  border-bottom-style: solid;
  border-bottom-width: 2px;
  border-bottom-color: #D3D3D3;
  vertical-align: bottom;
  padding-top: 5px;
  padding-bottom: 5px;
  overflow-x: hidden;
  display: inline-block;
  width: 100%;
}

.gt-post-table .gt_spanner_row {
  border-bottom-style: hidden;
}

.gt-post-table .gt_group_heading {
  padding-top: 8px;
  padding-bottom: 8px;
  padding-left: 5px;
  padding-right: 5px;
  color: #333333;
  background-color: #FFFFFF;
  font-size: 100%;
  font-weight: initial;
  text-transform: inherit;
  border-top-style: solid;
  border-top-width: 2px;
  border-top-color: #D3D3D3;
  border-bottom-style: solid;
  border-bottom-width: 2px;
  border-bottom-color: #D3D3D3;
  border-left-style: none;
  border-left-width: 1px;
  border-left-color: #D3D3D3;
  border-right-style: none;
  border-right-width: 1px;
  border-right-color: #D3D3D3;
  vertical-align: middle;
  text-align: left;
}

.gt-post-table .gt_empty_group_heading {
  padding: 0.5px;
  color: #333333;
  background-color: #FFFFFF;
  font-size: 100%;
  font-weight: initial;
  border-top-style: solid;
  border-top-width: 2px;
  border-top-color: #D3D3D3;
  border-bottom-style: solid;
  border-bottom-width: 2px;
  border-bottom-color: #D3D3D3;
  vertical-align: middle;
}

.gt-post-table .gt_from_md > :first-child {
  margin-top: 0;
}

.gt-post-table .gt_from_md > :last-child {
  margin-bottom: 0;
}

.gt-post-table .gt_row {
  padding-top: 8px;
  padding-bottom: 8px;
  padding-left: 5px;
  padding-right: 5px;
  margin: 10px;
  border-top-style: solid;
  border-top-width: 1px;
  border-top-color: #D3D3D3;
  border-left-style: none;
  border-left-width: 1px;
  border-left-color: #D3D3D3;
  border-right-style: none;
  border-right-width: 1px;
  border-right-color: #D3D3D3;
  vertical-align: middle;
  overflow-x: hidden;
}

.gt-post-table .gt_stub {
  color: #333333;
  background-color: #FFFFFF;
  font-size: 100%;
  font-weight: initial;
  text-transform: inherit;
  border-right-style: solid;
  border-right-width: 2px;
  border-right-color: #D3D3D3;
  padding-left: 5px;
  padding-right: 5px;
}

.gt-post-table .gt_stub_row_group {
  color: #333333;
  background-color: #FFFFFF;
  font-size: 100%;
  font-weight: initial;
  text-transform: inherit;
  border-right-style: solid;
  border-right-width: 2px;
  border-right-color: #D3D3D3;
  padding-left: 5px;
  padding-right: 5px;
  vertical-align: top;
}

.gt-post-table .gt_row_group_first td {
  border-top-width: 2px;
}

.gt-post-table .gt_row_group_first th {
  border-top-width: 2px;
}

.gt-post-table .gt_summary_row {
  color: #333333;
  background-color: #FFFFFF;
  text-transform: inherit;
  padding-top: 8px;
  padding-bottom: 8px;
  padding-left: 5px;
  padding-right: 5px;
}

.gt-post-table .gt_first_summary_row {
  border-top-style: solid;
  border-top-color: #D3D3D3;
}

.gt-post-table .gt_first_summary_row.thick {
  border-top-width: 2px;
}

.gt-post-table .gt_last_summary_row {
  padding-top: 8px;
  padding-bottom: 8px;
  padding-left: 5px;
  padding-right: 5px;
  border-bottom-style: solid;
  border-bottom-width: 2px;
  border-bottom-color: #D3D3D3;
}

.gt-post-table .gt_grand_summary_row {
  color: #333333;
  background-color: #FFFFFF;
  text-transform: inherit;
  padding-top: 8px;
  padding-bottom: 8px;
  padding-left: 5px;
  padding-right: 5px;
}

.gt-post-table .gt_first_grand_summary_row {
  padding-top: 8px;
  padding-bottom: 8px;
  padding-left: 5px;
  padding-right: 5px;
  border-top-style: double;
  border-top-width: 6px;
  border-top-color: #D3D3D3;
}

.gt-post-table .gt_last_grand_summary_row_top {
  padding-top: 8px;
  padding-bottom: 8px;
  padding-left: 5px;
  padding-right: 5px;
  border-bottom-style: double;
  border-bottom-width: 6px;
  border-bottom-color: #D3D3D3;
}

.gt-post-table .gt_striped {
  background-color: #F8F9FA;
}

.gt-post-table .gt_table_body {
  border-top-style: solid;
  border-top-width: 2px;
  border-top-color: #D3D3D3;
  border-bottom-style: solid;
  border-bottom-width: 2px;
  border-bottom-color: #D3D3D3;
}

.gt-post-table .gt_footnotes {
  color: #333333;
  background-color: #FFFFFF;
  border-bottom-style: none;
  border-bottom-width: 2px;
  border-bottom-color: #D3D3D3;
  border-left-style: none;
  border-left-width: 2px;
  border-left-color: #D3D3D3;
  border-right-style: none;
  border-right-width: 2px;
  border-right-color: #D3D3D3;
}

.gt-post-table .gt_footnote {
  margin: 0px;
  font-size: 90%;
  padding-top: 4px;
  padding-bottom: 4px;
  padding-left: 5px;
  padding-right: 5px;
}

.gt-post-table .gt_sourcenotes {
  color: #333333;
  background-color: #FFFFFF;
  border-bottom-style: none;
  border-bottom-width: 2px;
  border-bottom-color: #D3D3D3;
  border-left-style: none;
  border-left-width: 2px;
  border-left-color: #D3D3D3;
  border-right-style: none;
  border-right-width: 2px;
  border-right-color: #D3D3D3;
}

.gt-post-table .gt_sourcenote {
  font-size: 90%;
  padding-top: 4px;
  padding-bottom: 4px;
  padding-left: 5px;
  padding-right: 5px;
}

.gt-post-table .gt_left {
  text-align: left;
}

.gt-post-table .gt_center {
  text-align: center;
}

.gt-post-table .gt_right {
  text-align: right;
  font-variant-numeric: tabular-nums;
}

.gt-post-table .gt_font_normal {
  font-weight: normal;
}

.gt-post-table .gt_font_bold {
  font-weight: bold;
}

.gt-post-table .gt_font_italic {
  font-style: italic;
}

.gt-post-table .gt_super {
  font-size: 65%;
}

.gt-post-table .gt_footnote_marks {
  font-size: 75%;
  vertical-align: 0.4em;
  position: initial;
}

.gt-post-table .gt_asterisk {
  font-size: 100%;
  vertical-align: 0;
}

.gt-post-table .gt_indent_1 {
  text-indent: 5px;
}

.gt-post-table .gt_indent_2 {
  text-indent: 10px;
}

.gt-post-table .gt_indent_3 {
  text-indent: 15px;
}

.gt-post-table .gt_indent_4 {
  text-indent: 20px;
}

.gt-post-table .gt_indent_5 {
  text-indent: 25px;
}

.gt-post-table .katex-display {
  display: inline-flex !important;
  margin-bottom: 0.75em !important;
}

.gt-post-table div.Reactable > div.rt-table > div.rt-thead > div.rt-tr.rt-tr-group-header > div.rt-th-group:after {
  height: 0px !important;
}
</style>

<h2 id="the-reporting-problem">The Reporting Problem</h2>

<p>Machine learning in R is powerful, but reporting the results often takes
more effort than building the model itself. Every package returns
results in a different format — a named vector here, a matrix there, a
custom S3 object somewhere else. Even with excellent tidying packages
like <code class="language-plaintext highlighter-rouge">broom</code>, the column names and available fields change between model
types, and many ML packages aren’t supported at all.</p>

<p>The result is that producing consistent, polished visualisations and
tables for a report means writing custom extraction and reshaping code
for each package you use. Change your model, and your reporting code
breaks.</p>

<p><strong>tidylearn</strong> solves this. Every function — across 20 algorithms —
returns tidy tibbles, ggplot2 plots, and formatted <code class="language-plaintext highlighter-rouge">gt</code> tables with a
consistent structure. Your reporting pipeline becomes model-agnostic:
swap the algorithm, and the same plot code, table code, and comparison
logic work without modification.</p>

<p>This post walks through real-world analysis tasks — PCA, hierarchical
clustering, regularisation, and multi-model classification — comparing
the tidylearn workflow against the traditional approach.</p>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">library</span><span class="p">(</span><span class="n">tidylearn</span><span class="p">)</span><span class="w">
</span><span class="n">library</span><span class="p">(</span><span class="n">dplyr</span><span class="p">)</span><span class="w">
</span><span class="n">library</span><span class="p">(</span><span class="n">ggplot2</span><span class="p">)</span><span class="w">
</span><span class="n">library</span><span class="p">(</span><span class="n">gt</span><span class="p">)</span><span class="w">
</span><span class="n">library</span><span class="p">(</span><span class="n">tibble</span><span class="p">)</span><span class="w">
</span></code></pre></div></div>

<hr />

<h2 id="1-pca-biplot-scree-plot-and-variance-tables">1. PCA: Biplot, Scree Plot, and Variance Tables</h2>

<p>PCA is a staple of exploratory analysis, and producing a polished biplot
or scree plot is one of those tasks that should be straightforward but
rarely is.</p>

<h3 id="with-tidylearn">With tidylearn</h3>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">pca</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">tidy_pca</span><span class="p">(</span><span class="n">USArrests</span><span class="p">,</span><span class="w"> </span><span class="n">scale</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="kc">TRUE</span><span class="p">)</span><span class="w">
</span></code></pre></div></div>

<p>Scree plot with cumulative variance line and 80% threshold:</p>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">tidy_pca_screeplot</span><span class="p">(</span><span class="n">pca</span><span class="p">)</span><span class="w">
</span></code></pre></div></div>

<p><img src="https://blog.sheetsolved.com/assets/tidylearn/tl-pca-scree-1.png" alt="" /></p>

<p>Publication-ready biplot with observation scores and variable loadings:</p>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">tidy_pca_biplot</span><span class="p">(</span><span class="n">pca</span><span class="p">,</span><span class="w"> </span><span class="n">label_obs</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="kc">TRUE</span><span class="p">)</span><span class="w">
</span></code></pre></div></div>

<p><img src="https://blog.sheetsolved.com/assets/tidylearn/tl-pca-biplot-1.png" alt="" /></p>

<p>And the tables — variance explained and loadings — are one call each,
with colour-coded formatting out of the box:</p>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">pca_model</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">tl_model</span><span class="p">(</span><span class="n">USArrests</span><span class="p">,</span><span class="w"> </span><span class="n">method</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"pca"</span><span class="p">)</span><span class="w">
</span><span class="n">tl_table_variance</span><span class="p">(</span><span class="n">pca_model</span><span class="p">)</span><span class="w">
</span></code></pre></div></div>

<div id="qeayjdrzqo" class="gt-post-table" style="padding-left:0px;padding-right:0px;padding-top:10px;padding-bottom:10px;overflow-x:auto;overflow-y:auto;width:auto;height:auto;">

<table class="gt_table" data-quarto-disable-processing="false" data-quarto-bootstrap="false">
  <thead>
    <tr class="gt_heading">
      <td colspan="5" class="gt_heading gt_title gt_font_normal gt_bottom_border" style="">PCA Variance Explained</td>
    </tr>
    
    <tr class="gt_col_headings">
      <th class="gt_col_heading gt_columns_bottom_border gt_left" rowspan="1" colspan="1" style="color: #FFFFFF;" scope="col" id="component">Component</th>
      <th class="gt_col_heading gt_columns_bottom_border gt_right" rowspan="1" colspan="1" style="color: #FFFFFF;" scope="col" id="sdev">Std. Dev.</th>
      <th class="gt_col_heading gt_columns_bottom_border gt_right" rowspan="1" colspan="1" style="color: #FFFFFF;" scope="col" id="variance">Variance</th>
      <th class="gt_col_heading gt_columns_bottom_border gt_right" rowspan="1" colspan="1" style="color: #FFFFFF;" scope="col" id="prop_variance">Proportion</th>
      <th class="gt_col_heading gt_columns_bottom_border gt_right" rowspan="1" colspan="1" style="color: #FFFFFF;" scope="col" id="cum_variance">Cumulative</th>
    </tr>
  </thead>
  <tbody class="gt_table_body">
    <tr><td headers="component" class="gt_row gt_left">PC1</td>
<td headers="sdev" class="gt_row gt_right">1.5749</td>
<td headers="variance" class="gt_row gt_right">2.4802</td>
<td headers="prop_variance" class="gt_row gt_right">62.0%</td>
<td headers="cum_variance" class="gt_row gt_right" style="background-color: #FFFFFF; color: #000000;">62.0%</td></tr>
    <tr><td headers="component" class="gt_row gt_left gt_striped">PC2</td>
<td headers="sdev" class="gt_row gt_right gt_striped">0.9949</td>
<td headers="variance" class="gt_row gt_right gt_striped">0.9898</td>
<td headers="prop_variance" class="gt_row gt_right gt_striped">24.7%</td>
<td headers="cum_variance" class="gt_row gt_right gt_striped" style="background-color: #81CB96; color: #000000;">86.8%</td></tr>
    <tr><td headers="component" class="gt_row gt_left">PC3</td>
<td headers="sdev" class="gt_row gt_right">0.5971</td>
<td headers="variance" class="gt_row gt_right">0.3566</td>
<td headers="prop_variance" class="gt_row gt_right">8.9%</td>
<td headers="cum_variance" class="gt_row gt_right" style="background-color: #4CB871; color: #000000;">95.7%</td></tr>
    <tr><td headers="component" class="gt_row gt_left gt_striped">PC4</td>
<td headers="sdev" class="gt_row gt_right gt_striped">0.4164</td>
<td headers="variance" class="gt_row gt_right gt_striped">0.1734</td>
<td headers="prop_variance" class="gt_row gt_right gt_striped">4.3%</td>
<td headers="cum_variance" class="gt_row gt_right gt_striped" style="background-color: #27AE60; color: #FFFFFF;">100.0%</td></tr>
  </tbody>
  <tfoot>
    <tr class="gt_sourcenotes">
      <td class="gt_sourcenote" colspan="5">tidylearn | pca | n = 50</td>
    </tr>
  </tfoot>
</table>
</div>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">tl_table_loadings</span><span class="p">(</span><span class="n">pca_model</span><span class="p">,</span><span class="w"> </span><span class="n">n_components</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="m">2</span><span class="p">)</span><span class="w">
</span></code></pre></div></div>

<div id="xgdyentzro" class="gt-post-table" style="padding-left:0px;padding-right:0px;padding-top:10px;padding-bottom:10px;overflow-x:auto;overflow-y:auto;width:auto;height:auto;">

<table class="gt_table" data-quarto-disable-processing="false" data-quarto-bootstrap="false">
  <thead>
    <tr class="gt_heading">
      <td colspan="3" class="gt_heading gt_title gt_font_normal gt_bottom_border" style="">PCA Loadings</td>
    </tr>
    
    <tr class="gt_col_headings">
      <th class="gt_col_heading gt_columns_bottom_border gt_left" rowspan="1" colspan="1" style="color: #FFFFFF;" scope="col" id="variable">Variable</th>
      <th class="gt_col_heading gt_columns_bottom_border gt_right" rowspan="1" colspan="1" style="color: #FFFFFF;" scope="col" id="PC1">PC1</th>
      <th class="gt_col_heading gt_columns_bottom_border gt_right" rowspan="1" colspan="1" style="color: #FFFFFF;" scope="col" id="PC2">PC2</th>
    </tr>
  </thead>
  <tbody class="gt_table_body">
    <tr><td headers="variable" class="gt_row gt_left">Murder</td>
<td headers="PC1" class="gt_row gt_right" style="background-color: #E89887; color: #000000;">−0.536</td>
<td headers="PC2" class="gt_row gt_right" style="background-color: #EFAEA1; color: #000000;">−0.418</td></tr>
    <tr><td headers="variable" class="gt_row gt_left gt_striped">Assault</td>
<td headers="PC1" class="gt_row gt_right gt_striped" style="background-color: #E58E7D; color: #000000;">−0.583</td>
<td headers="PC2" class="gt_row gt_right gt_striped" style="background-color: #FADAD4; color: #000000;">−0.188</td></tr>
    <tr><td headers="variable" class="gt_row gt_left">UrbanPop</td>
<td headers="PC1" class="gt_row gt_right" style="background-color: #F7C9BF; color: #000000;">−0.278</td>
<td headers="PC2" class="gt_row gt_right" style="background-color: #518FC2; color: #FFFFFF;">0.873</td></tr>
    <tr><td headers="variable" class="gt_row gt_left gt_striped">Rape</td>
<td headers="PC1" class="gt_row gt_right gt_striped" style="background-color: #E89686; color: #000000;">−0.543</td>
<td headers="PC2" class="gt_row gt_right gt_striped" style="background-color: #E0E9F3; color: #000000;">0.167</td></tr>
  </tbody>
  <tfoot>
    <tr class="gt_sourcenotes">
      <td class="gt_sourcenote" colspan="3">tidylearn | pca | n = 50</td>
    </tr>
  </tfoot>
</table>
</div>

<p>The loadings table uses a diverging red–blue colour scale to highlight
strong positive and negative loadings — no manual formatting required.</p>

<h3 id="without-tidylearn">Without tidylearn</h3>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">pca_base</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">prcomp</span><span class="p">(</span><span class="n">USArrests</span><span class="p">,</span><span class="w"> </span><span class="n">scale.</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="kc">TRUE</span><span class="p">)</span><span class="w">
</span></code></pre></div></div>

<p>The default R biplot:</p>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">biplot</span><span class="p">(</span><span class="n">pca_base</span><span class="p">)</span><span class="w">
</span></code></pre></div></div>

<p><img src="https://blog.sheetsolved.com/assets/tidylearn/base-pca-biplot-1.png" alt="" /></p>

<p>This produces a functional but visually rough base R graphic — no
<code class="language-plaintext highlighter-rouge">theme_minimal()</code>, no consistent colour scheme, no variance-explained
axis labels. To get a ggplot2 biplot, you need to manually extract and
scale the scores and loadings:</p>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># Extract scores</span><span class="w">
</span><span class="n">scores</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">as.data.frame</span><span class="p">(</span><span class="n">pca_base</span><span class="o">$</span><span class="n">x</span><span class="p">[,</span><span class="w"> </span><span class="m">1</span><span class="o">:</span><span class="m">2</span><span class="p">])</span><span class="w">
</span><span class="n">scores</span><span class="o">$</span><span class="n">label</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">rownames</span><span class="p">(</span><span class="n">scores</span><span class="p">)</span><span class="w">

</span><span class="c1"># Extract loadings and scale to match scores</span><span class="w">
</span><span class="n">loadings</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">as.data.frame</span><span class="p">(</span><span class="n">pca_base</span><span class="o">$</span><span class="n">rotation</span><span class="p">[,</span><span class="w"> </span><span class="m">1</span><span class="o">:</span><span class="m">2</span><span class="p">])</span><span class="w">
</span><span class="n">loadings</span><span class="o">$</span><span class="n">variable</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">rownames</span><span class="p">(</span><span class="n">loadings</span><span class="p">)</span><span class="w">
</span><span class="n">score_range</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="nf">max</span><span class="p">(</span><span class="nf">abs</span><span class="p">(</span><span class="n">scores</span><span class="p">[,</span><span class="w"> </span><span class="m">1</span><span class="o">:</span><span class="m">2</span><span class="p">]))</span><span class="w">
</span><span class="n">loading_range</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="nf">max</span><span class="p">(</span><span class="nf">abs</span><span class="p">(</span><span class="n">loadings</span><span class="p">[,</span><span class="w"> </span><span class="m">1</span><span class="o">:</span><span class="m">2</span><span class="p">]))</span><span class="w">
</span><span class="n">scale_factor</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="p">(</span><span class="n">score_range</span><span class="w"> </span><span class="o">/</span><span class="w"> </span><span class="n">loading_range</span><span class="p">)</span><span class="w"> </span><span class="o">*</span><span class="w"> </span><span class="m">0.8</span><span class="w">

</span><span class="n">loadings</span><span class="o">$</span><span class="n">PC1_scaled</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">loadings</span><span class="o">$</span><span class="n">PC1</span><span class="w"> </span><span class="o">*</span><span class="w"> </span><span class="n">scale_factor</span><span class="w">
</span><span class="n">loadings</span><span class="o">$</span><span class="n">PC2_scaled</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">loadings</span><span class="o">$</span><span class="n">PC2</span><span class="w"> </span><span class="o">*</span><span class="w"> </span><span class="n">scale_factor</span><span class="w">

</span><span class="c1"># Compute variance explained for axis labels</span><span class="w">
</span><span class="n">var_exp</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">pca_base</span><span class="o">$</span><span class="n">sdev</span><span class="o">^</span><span class="m">2</span><span class="w"> </span><span class="o">/</span><span class="w"> </span><span class="nf">sum</span><span class="p">(</span><span class="n">pca_base</span><span class="o">$</span><span class="n">sdev</span><span class="o">^</span><span class="m">2</span><span class="p">)</span><span class="w"> </span><span class="o">*</span><span class="w"> </span><span class="m">100</span><span class="w">

</span><span class="c1"># Build the ggplot manually</span><span class="w">
</span><span class="n">ggplot</span><span class="p">()</span><span class="w"> </span><span class="o">+</span><span class="w">
  </span><span class="n">geom_point</span><span class="p">(</span><span class="n">data</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">scores</span><span class="p">,</span><span class="w"> </span><span class="n">aes</span><span class="p">(</span><span class="n">x</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PC1</span><span class="p">,</span><span class="w"> </span><span class="n">y</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PC2</span><span class="p">),</span><span class="w">
             </span><span class="n">colour</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"steelblue"</span><span class="p">,</span><span class="w"> </span><span class="n">alpha</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="m">0.7</span><span class="p">)</span><span class="w"> </span><span class="o">+</span><span class="w">
  </span><span class="n">geom_text</span><span class="p">(</span><span class="n">data</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">scores</span><span class="p">,</span><span class="w"> </span><span class="n">aes</span><span class="p">(</span><span class="n">x</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PC1</span><span class="p">,</span><span class="w"> </span><span class="n">y</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PC2</span><span class="p">,</span><span class="w"> </span><span class="n">label</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">label</span><span class="p">),</span><span class="w">
            </span><span class="n">size</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="m">2.5</span><span class="p">,</span><span class="w"> </span><span class="n">vjust</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="m">-0.5</span><span class="p">)</span><span class="w"> </span><span class="o">+</span><span class="w">
  </span><span class="n">geom_segment</span><span class="p">(</span><span class="n">data</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">loadings</span><span class="p">,</span><span class="w">
               </span><span class="n">aes</span><span class="p">(</span><span class="n">x</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="m">0</span><span class="p">,</span><span class="w"> </span><span class="n">y</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="m">0</span><span class="p">,</span><span class="w"> </span><span class="n">xend</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PC1_scaled</span><span class="p">,</span><span class="w"> </span><span class="n">yend</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PC2_scaled</span><span class="p">),</span><span class="w">
               </span><span class="n">arrow</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">arrow</span><span class="p">(</span><span class="n">length</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">unit</span><span class="p">(</span><span class="m">0.2</span><span class="p">,</span><span class="w"> </span><span class="s2">"cm"</span><span class="p">)),</span><span class="w">
               </span><span class="n">colour</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"red"</span><span class="p">,</span><span class="w"> </span><span class="n">alpha</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="m">0.7</span><span class="p">)</span><span class="w"> </span><span class="o">+</span><span class="w">
  </span><span class="n">geom_text</span><span class="p">(</span><span class="n">data</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">loadings</span><span class="p">,</span><span class="w">
            </span><span class="n">aes</span><span class="p">(</span><span class="n">x</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PC1_scaled</span><span class="p">,</span><span class="w"> </span><span class="n">y</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PC2_scaled</span><span class="p">,</span><span class="w"> </span><span class="n">label</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">variable</span><span class="p">),</span><span class="w">
            </span><span class="n">colour</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"red"</span><span class="p">,</span><span class="w"> </span><span class="n">size</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="m">3</span><span class="p">,</span><span class="w"> </span><span class="n">fontface</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"bold"</span><span class="p">,</span><span class="w"> </span><span class="n">vjust</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="m">-0.5</span><span class="p">)</span><span class="w"> </span><span class="o">+</span><span class="w">
  </span><span class="n">labs</span><span class="p">(</span><span class="w">
    </span><span class="n">title</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"PCA Biplot"</span><span class="p">,</span><span class="w">
    </span><span class="n">x</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">sprintf</span><span class="p">(</span><span class="s2">"PC1 (%.1f%% variance)"</span><span class="p">,</span><span class="w"> </span><span class="n">var_exp</span><span class="p">[</span><span class="m">1</span><span class="p">]),</span><span class="w">
    </span><span class="n">y</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">sprintf</span><span class="p">(</span><span class="s2">"PC2 (%.1f%% variance)"</span><span class="p">,</span><span class="w"> </span><span class="n">var_exp</span><span class="p">[</span><span class="m">2</span><span class="p">])</span><span class="w">
  </span><span class="p">)</span><span class="w"> </span><span class="o">+</span><span class="w">
  </span><span class="n">coord_equal</span><span class="p">()</span><span class="w"> </span><span class="o">+</span><span class="w">
  </span><span class="n">theme_minimal</span><span class="p">()</span><span class="w">
</span></code></pre></div></div>

<p><img src="https://blog.sheetsolved.com/assets/tidylearn/base-pca-ggbiplot-1.png" alt="" /></p>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># Variance table — manually computed</span><span class="w">
</span><span class="n">var_explained</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">data.frame</span><span class="p">(</span><span class="w">
  </span><span class="n">component</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">paste0</span><span class="p">(</span><span class="s2">"PC"</span><span class="p">,</span><span class="w"> </span><span class="nf">seq_along</span><span class="p">(</span><span class="n">pca_base</span><span class="o">$</span><span class="n">sdev</span><span class="p">)),</span><span class="w">
  </span><span class="n">sdev</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">pca_base</span><span class="o">$</span><span class="n">sdev</span><span class="p">,</span><span class="w">
  </span><span class="n">variance</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">pca_base</span><span class="o">$</span><span class="n">sdev</span><span class="o">^</span><span class="m">2</span><span class="p">,</span><span class="w">
  </span><span class="n">prop_variance</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">pca_base</span><span class="o">$</span><span class="n">sdev</span><span class="o">^</span><span class="m">2</span><span class="w"> </span><span class="o">/</span><span class="w"> </span><span class="nf">sum</span><span class="p">(</span><span class="n">pca_base</span><span class="o">$</span><span class="n">sdev</span><span class="o">^</span><span class="m">2</span><span class="p">),</span><span class="w">
  </span><span class="n">cum_variance</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="nf">cumsum</span><span class="p">(</span><span class="n">pca_base</span><span class="o">$</span><span class="n">sdev</span><span class="o">^</span><span class="m">2</span><span class="w"> </span><span class="o">/</span><span class="w"> </span><span class="nf">sum</span><span class="p">(</span><span class="n">pca_base</span><span class="o">$</span><span class="n">sdev</span><span class="o">^</span><span class="m">2</span><span class="p">))</span><span class="w">
</span><span class="p">)</span><span class="w">

</span><span class="n">knitr</span><span class="o">::</span><span class="n">kable</span><span class="p">(</span><span class="n">var_explained</span><span class="p">,</span><span class="w"> </span><span class="n">digits</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="m">3</span><span class="p">,</span><span class="w">
             </span><span class="n">caption</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"Variance Explained (manual)"</span><span class="p">)</span><span class="w">
</span></code></pre></div></div>

<table>
  <thead>
    <tr>
      <th style="text-align: left">component</th>
      <th style="text-align: right">sdev</th>
      <th style="text-align: right">variance</th>
      <th style="text-align: right">prop_variance</th>
      <th style="text-align: right">cum_variance</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left">PC1</td>
      <td style="text-align: right">1.575</td>
      <td style="text-align: right">2.480</td>
      <td style="text-align: right">0.620</td>
      <td style="text-align: right">0.620</td>
    </tr>
    <tr>
      <td style="text-align: left">PC2</td>
      <td style="text-align: right">0.995</td>
      <td style="text-align: right">0.990</td>
      <td style="text-align: right">0.247</td>
      <td style="text-align: right">0.868</td>
    </tr>
    <tr>
      <td style="text-align: left">PC3</td>
      <td style="text-align: right">0.597</td>
      <td style="text-align: right">0.357</td>
      <td style="text-align: right">0.089</td>
      <td style="text-align: right">0.957</td>
    </tr>
    <tr>
      <td style="text-align: left">PC4</td>
      <td style="text-align: right">0.416</td>
      <td style="text-align: right">0.173</td>
      <td style="text-align: right">0.043</td>
      <td style="text-align: right">1.000</td>
    </tr>
  </tbody>
</table>

<p>Variance Explained (manual)</p>

<p>The manual biplot requires extracting matrices, scaling loadings to
match score ranges, computing variance percentages for axis labels, and
assembling the ggplot layer by layer. <code class="language-plaintext highlighter-rouge">tidy_pca_biplot()</code> handles all of
that in one call. And the <code class="language-plaintext highlighter-rouge">kable()</code> variance table is functional but
plain — no colour coding, no cumulative-variance highlighting.
<code class="language-plaintext highlighter-rouge">tl_table_variance()</code> adds these by default.</p>

<hr />

<h2 id="2-hierarchical-clustering">2. Hierarchical Clustering</h2>

<p>Hierarchical clustering involves computing distances, fitting the tree,
visualising the dendrogram, cutting it, and augmenting your data with
cluster assignments. Each step traditionally produces a different data
structure.</p>

<h3 id="with-tidylearn-1">With tidylearn</h3>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">hc</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">tidy_hclust</span><span class="p">(</span><span class="n">USArrests</span><span class="p">,</span><span class="w"> </span><span class="n">method</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"ward.D2"</span><span class="p">)</span><span class="w">

</span><span class="c1"># Dendrogram with cluster rectangles</span><span class="w">
</span><span class="n">tidy_dendrogram</span><span class="p">(</span><span class="n">hc</span><span class="p">,</span><span class="w"> </span><span class="n">k</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="m">4</span><span class="p">)</span><span class="w">
</span></code></pre></div></div>

<p><img src="https://blog.sheetsolved.com/assets/tidylearn/tl-hclust-1.png" alt="" /></p>

<p>Cut the tree and get a tidy tibble of cluster assignments:</p>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">clusters</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">tidy_cutree</span><span class="p">(</span><span class="n">hc</span><span class="p">,</span><span class="w"> </span><span class="n">k</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="m">4</span><span class="p">)</span><span class="w">
</span><span class="n">knitr</span><span class="o">::</span><span class="n">kable</span><span class="p">(</span><span class="n">head</span><span class="p">(</span><span class="n">clusters</span><span class="p">,</span><span class="w"> </span><span class="m">10</span><span class="p">),</span><span class="w"> </span><span class="n">caption</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"Cluster Assignments (first 10)"</span><span class="p">)</span><span class="w">
</span></code></pre></div></div>

<table>
  <thead>
    <tr>
      <th style="text-align: left">.obs_id</th>
      <th style="text-align: right">cluster</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left">Alabama</td>
      <td style="text-align: right">1</td>
    </tr>
    <tr>
      <td style="text-align: left">Alaska</td>
      <td style="text-align: right">1</td>
    </tr>
    <tr>
      <td style="text-align: left">Arizona</td>
      <td style="text-align: right">1</td>
    </tr>
    <tr>
      <td style="text-align: left">Arkansas</td>
      <td style="text-align: right">2</td>
    </tr>
    <tr>
      <td style="text-align: left">California</td>
      <td style="text-align: right">1</td>
    </tr>
    <tr>
      <td style="text-align: left">Colorado</td>
      <td style="text-align: right">2</td>
    </tr>
    <tr>
      <td style="text-align: left">Connecticut</td>
      <td style="text-align: right">3</td>
    </tr>
    <tr>
      <td style="text-align: left">Delaware</td>
      <td style="text-align: right">1</td>
    </tr>
    <tr>
      <td style="text-align: left">Florida</td>
      <td style="text-align: right">1</td>
    </tr>
    <tr>
      <td style="text-align: left">Georgia</td>
      <td style="text-align: right">2</td>
    </tr>
  </tbody>
</table>

<p>Cluster Assignments (first 10)</p>

<p>Augment the original data and produce a formatted cluster summary table:</p>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">hc_model</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">tl_model</span><span class="p">(</span><span class="n">USArrests</span><span class="p">,</span><span class="w"> </span><span class="n">method</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"hclust"</span><span class="p">)</span><span class="w">
</span><span class="n">tl_table_clusters</span><span class="p">(</span><span class="n">hc_model</span><span class="p">,</span><span class="w"> </span><span class="n">k</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="m">4</span><span class="p">)</span><span class="w">
</span></code></pre></div></div>

<div id="ciydemetbh" class="gt-post-table" style="padding-left:0px;padding-right:0px;padding-top:10px;padding-bottom:10px;overflow-x:auto;overflow-y:auto;width:auto;height:auto;">

<table class="gt_table" data-quarto-disable-processing="false" data-quarto-bootstrap="false">
  <thead>
    <tr class="gt_heading">
      <td colspan="6" class="gt_heading gt_title gt_font_normal" style="">Cluster Summary</td>
    </tr>
    <tr class="gt_heading">
      <td colspan="6" class="gt_heading gt_subtitle gt_font_normal gt_bottom_border" style="">hclust | 4 clusters</td>
    </tr>
    <tr class="gt_col_headings">
      <th class="gt_col_heading gt_columns_bottom_border gt_right" rowspan="1" colspan="1" style="color: #FFFFFF;" scope="col" id="cluster">Cluster</th>
      <th class="gt_col_heading gt_columns_bottom_border gt_right" rowspan="1" colspan="1" style="color: #FFFFFF;" scope="col" id="size">Size</th>
      <th class="gt_col_heading gt_columns_bottom_border gt_right" rowspan="1" colspan="1" style="color: #FFFFFF;" scope="col" id="Murder">Murder</th>
      <th class="gt_col_heading gt_columns_bottom_border gt_right" rowspan="1" colspan="1" style="color: #FFFFFF;" scope="col" id="Assault">Assault</th>
      <th class="gt_col_heading gt_columns_bottom_border gt_right" rowspan="1" colspan="1" style="color: #FFFFFF;" scope="col" id="UrbanPop">UrbanPop</th>
      <th class="gt_col_heading gt_columns_bottom_border gt_right" rowspan="1" colspan="1" style="color: #FFFFFF;" scope="col" id="Rape">Rape</th>
    </tr>
  </thead>
  <tbody class="gt_table_body">
    <tr><td headers="cluster" class="gt_row gt_right">1</td>
<td headers="size" class="gt_row gt_right">14</td>
<td headers="Murder" class="gt_row gt_right">11.47</td>
<td headers="Assault" class="gt_row gt_right">263.50</td>
<td headers="UrbanPop" class="gt_row gt_right">69.14</td>
<td headers="Rape" class="gt_row gt_right">29.00</td></tr>
    <tr><td headers="cluster" class="gt_row gt_right gt_striped">2</td>
<td headers="size" class="gt_row gt_right gt_striped">14</td>
<td headers="Murder" class="gt_row gt_right gt_striped">8.21</td>
<td headers="Assault" class="gt_row gt_right gt_striped">173.29</td>
<td headers="UrbanPop" class="gt_row gt_right gt_striped">70.64</td>
<td headers="Rape" class="gt_row gt_right gt_striped">22.84</td></tr>
    <tr><td headers="cluster" class="gt_row gt_right">3</td>
<td headers="size" class="gt_row gt_right">20</td>
<td headers="Murder" class="gt_row gt_right">4.27</td>
<td headers="Assault" class="gt_row gt_right">87.55</td>
<td headers="UrbanPop" class="gt_row gt_right">59.75</td>
<td headers="Rape" class="gt_row gt_right">14.39</td></tr>
    <tr><td headers="cluster" class="gt_row gt_right gt_striped">4</td>
<td headers="size" class="gt_row gt_right gt_striped">2</td>
<td headers="Murder" class="gt_row gt_right gt_striped">14.20</td>
<td headers="Assault" class="gt_row gt_right gt_striped">336.00</td>
<td headers="UrbanPop" class="gt_row gt_right gt_striped">62.50</td>
<td headers="Rape" class="gt_row gt_right gt_striped">24.00</td></tr>
  </tbody>
  <tfoot>
    <tr class="gt_sourcenotes">
      <td class="gt_sourcenote" colspan="6">tidylearn | hclust | n = 50</td>
    </tr>
  </tfoot>
</table>
</div>

<h3 id="without-tidylearn-1">Without tidylearn</h3>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># Compute distance matrix</span><span class="w">
</span><span class="n">d</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">dist</span><span class="p">(</span><span class="n">scale</span><span class="p">(</span><span class="n">USArrests</span><span class="p">),</span><span class="w"> </span><span class="n">method</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"euclidean"</span><span class="p">)</span><span class="w">

</span><span class="c1"># Fit hierarchical clustering</span><span class="w">
</span><span class="n">hc_base</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">hclust</span><span class="p">(</span><span class="n">d</span><span class="p">,</span><span class="w"> </span><span class="n">method</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"ward.D2"</span><span class="p">)</span><span class="w">

</span><span class="c1"># Plot dendrogram</span><span class="w">
</span><span class="n">plot</span><span class="p">(</span><span class="n">hc_base</span><span class="p">,</span><span class="w"> </span><span class="n">main</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"Hierarchical Clustering Dendrogram"</span><span class="p">,</span><span class="w">
     </span><span class="n">xlab</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">""</span><span class="p">,</span><span class="w"> </span><span class="n">sub</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">""</span><span class="p">,</span><span class="w"> </span><span class="n">cex</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="m">0.7</span><span class="p">)</span><span class="w">
</span><span class="n">rect.hclust</span><span class="p">(</span><span class="n">hc_base</span><span class="p">,</span><span class="w"> </span><span class="n">k</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="m">4</span><span class="p">,</span><span class="w"> </span><span class="n">border</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="m">2</span><span class="o">:</span><span class="m">5</span><span class="p">)</span><span class="w">
</span></code></pre></div></div>

<p><img src="https://blog.sheetsolved.com/assets/tidylearn/base-hclust-1.png" alt="" /></p>

<p>The dendrogram looks similar — both use base R graphics for this. The
real difference is in what happens next:</p>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># cutree returns a named integer vector — not a tibble</span><span class="w">
</span><span class="n">clusters_base</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">cutree</span><span class="p">(</span><span class="n">hc_base</span><span class="p">,</span><span class="w"> </span><span class="n">k</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="m">4</span><span class="p">)</span><span class="w">
</span><span class="n">str</span><span class="p">(</span><span class="n">clusters_base</span><span class="p">)</span><span class="w">
</span><span class="c1">#&gt;  Named int [1:50] 1 2 2 3 2 2 3 3 2 1 ...</span><span class="w">
</span><span class="c1">#&gt;  - attr(*, "names")= chr [1:50] "Alabama" "Alaska" "Arizona" "Arkansas" ...</span><span class="w">
</span></code></pre></div></div>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># To get a summary table, manually bind and reshape</span><span class="w">
</span><span class="n">USArrests_clustered</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">USArrests</span><span class="w">
</span><span class="n">USArrests_clustered</span><span class="o">$</span><span class="n">cluster</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">clusters_base</span><span class="w">

</span><span class="n">USArrests_clustered</span><span class="w"> </span><span class="o">%&gt;%</span><span class="w">
  </span><span class="n">group_by</span><span class="p">(</span><span class="n">cluster</span><span class="p">)</span><span class="w"> </span><span class="o">%&gt;%</span><span class="w">
  </span><span class="n">summarise</span><span class="p">(</span><span class="n">across</span><span class="p">(</span><span class="n">where</span><span class="p">(</span><span class="n">is.numeric</span><span class="p">),</span><span class="w"> </span><span class="n">mean</span><span class="p">),</span><span class="w"> </span><span class="n">.groups</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"drop"</span><span class="p">)</span><span class="w"> </span><span class="o">%&gt;%</span><span class="w">
  </span><span class="n">knitr</span><span class="o">::</span><span class="n">kable</span><span class="p">(</span><span class="n">digits</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="m">1</span><span class="p">,</span><span class="w"> </span><span class="n">caption</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"Cluster Means (manual)"</span><span class="p">)</span><span class="w">
</span></code></pre></div></div>

<table>
  <thead>
    <tr>
      <th style="text-align: right">cluster</th>
      <th style="text-align: right">Murder</th>
      <th style="text-align: right">Assault</th>
      <th style="text-align: right">UrbanPop</th>
      <th style="text-align: right">Rape</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: right">1</td>
      <td style="text-align: right">14.7</td>
      <td style="text-align: right">251.3</td>
      <td style="text-align: right">54.3</td>
      <td style="text-align: right">21.7</td>
    </tr>
    <tr>
      <td style="text-align: right">2</td>
      <td style="text-align: right">11.0</td>
      <td style="text-align: right">264.0</td>
      <td style="text-align: right">76.5</td>
      <td style="text-align: right">33.6</td>
    </tr>
    <tr>
      <td style="text-align: right">3</td>
      <td style="text-align: right">6.2</td>
      <td style="text-align: right">142.1</td>
      <td style="text-align: right">71.3</td>
      <td style="text-align: right">19.2</td>
    </tr>
    <tr>
      <td style="text-align: right">4</td>
      <td style="text-align: right">3.1</td>
      <td style="text-align: right">76.0</td>
      <td style="text-align: right">52.1</td>
      <td style="text-align: right">11.8</td>
    </tr>
  </tbody>
</table>

<p>Cluster Means (manual)</p>

<p>The dendrogram itself is comparable. But <code class="language-plaintext highlighter-rouge">cutree()</code> returns a named
integer vector that needs manual binding to your data, and the resulting
<code class="language-plaintext highlighter-rouge">kable()</code> is plain text. <code class="language-plaintext highlighter-rouge">tl_table_clusters()</code> produces a formatted
table with cluster sizes, styled headers, and consistent theming — ready
for a report.</p>

<hr />

<h2 id="3-lasso-regularisation-coefficient-paths-and-tables">3. Lasso Regularisation: Coefficient Paths and Tables</h2>

<p>Regularised models are a common choice, but visualising how coefficients
shrink along the regularisation path is one of those tasks where the
default output is decidedly not report-ready.</p>

<h3 id="with-tidylearn-2">With tidylearn</h3>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">lasso</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">tl_model</span><span class="p">(</span><span class="n">mtcars</span><span class="p">,</span><span class="w"> </span><span class="n">mpg</span><span class="w"> </span><span class="o">~</span><span class="w"> </span><span class="n">.</span><span class="p">,</span><span class="w"> </span><span class="n">method</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"lasso"</span><span class="p">)</span><span class="w">

</span><span class="c1"># Coefficient path as a ggplot2 object</span><span class="w">
</span><span class="n">tl_plot_regularization_path</span><span class="p">(</span><span class="n">lasso</span><span class="p">)</span><span class="w">
</span></code></pre></div></div>

<p><img src="https://blog.sheetsolved.com/assets/tidylearn/tl-lasso-1.png" alt="" /></p>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># Cross-validation curve</span><span class="w">
</span><span class="n">tl_plot_regularization_cv</span><span class="p">(</span><span class="n">lasso</span><span class="p">)</span><span class="w">
</span></code></pre></div></div>

<p><img src="https://blog.sheetsolved.com/assets/tidylearn/tl-lasso-cv-1.png" alt="" /></p>

<p>And a formatted coefficient table, sorted by magnitude with zero
coefficients greyed out:</p>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">tl_table_coefficients</span><span class="p">(</span><span class="n">lasso</span><span class="p">)</span><span class="w">
</span></code></pre></div></div>

<div id="wfzjheaqgd" class="gt-post-table" style="padding-left:0px;padding-right:0px;padding-top:10px;padding-bottom:10px;overflow-x:auto;overflow-y:auto;width:auto;height:auto;">

<table class="gt_table" data-quarto-disable-processing="false" data-quarto-bootstrap="false">
  <thead>
    <tr class="gt_heading">
      <td colspan="3" class="gt_heading gt_title gt_font_normal" style="">Lasso Coefficients</td>
    </tr>
    <tr class="gt_heading">
      <td colspan="3" class="gt_heading gt_subtitle gt_font_normal gt_bottom_border" style="">lambda = 1.275 (1se)</td>
    </tr>
    <tr class="gt_col_headings">
      <th class="gt_col_heading gt_columns_bottom_border gt_left" rowspan="1" colspan="1" style="color: #FFFFFF;" scope="col" id="term">Term</th>
      <th class="gt_col_heading gt_columns_bottom_border gt_right" rowspan="1" colspan="1" style="color: #FFFFFF;" scope="col" id="estimate">Coefficient</th>
      <th class="gt_col_heading gt_columns_bottom_border gt_right" rowspan="1" colspan="1" style="color: #FFFFFF;" scope="col" id="abs_estimate">|Coefficient|</th>
    </tr>
  </thead>
  <tbody class="gt_table_body">
    <tr><td headers="term" class="gt_row gt_left">(Intercept)</td>
<td headers="estimate" class="gt_row gt_right">34.3695</td>
<td headers="abs_estimate" class="gt_row gt_right">34.3695</td></tr>
    <tr><td headers="term" class="gt_row gt_left gt_striped">wt</td>
<td headers="estimate" class="gt_row gt_right gt_striped">−2.4351</td>
<td headers="abs_estimate" class="gt_row gt_right gt_striped">2.4351</td></tr>
    <tr><td headers="term" class="gt_row gt_left">cyl</td>
<td headers="estimate" class="gt_row gt_right">−0.8528</td>
<td headers="abs_estimate" class="gt_row gt_right">0.8528</td></tr>
    <tr><td headers="term" class="gt_row gt_left gt_striped">hp</td>
<td headers="estimate" class="gt_row gt_right gt_striped">−0.0080</td>
<td headers="abs_estimate" class="gt_row gt_right gt_striped">0.0080</td></tr>
    <tr><td headers="term" class="gt_row gt_left" style="color: #999999;">disp</td>
<td headers="estimate" class="gt_row gt_right" style="color: #999999;">0.0000</td>
<td headers="abs_estimate" class="gt_row gt_right" style="color: #999999;">0.0000</td></tr>
    <tr><td headers="term" class="gt_row gt_left gt_striped" style="color: #999999;">drat</td>
<td headers="estimate" class="gt_row gt_right gt_striped" style="color: #999999;">0.0000</td>
<td headers="abs_estimate" class="gt_row gt_right gt_striped" style="color: #999999;">0.0000</td></tr>
    <tr><td headers="term" class="gt_row gt_left" style="color: #999999;">qsec</td>
<td headers="estimate" class="gt_row gt_right" style="color: #999999;">0.0000</td>
<td headers="abs_estimate" class="gt_row gt_right" style="color: #999999;">0.0000</td></tr>
    <tr><td headers="term" class="gt_row gt_left gt_striped" style="color: #999999;">vs</td>
<td headers="estimate" class="gt_row gt_right gt_striped" style="color: #999999;">0.0000</td>
<td headers="abs_estimate" class="gt_row gt_right gt_striped" style="color: #999999;">0.0000</td></tr>
    <tr><td headers="term" class="gt_row gt_left" style="color: #999999;">am</td>
<td headers="estimate" class="gt_row gt_right" style="color: #999999;">0.0000</td>
<td headers="abs_estimate" class="gt_row gt_right" style="color: #999999;">0.0000</td></tr>
    <tr><td headers="term" class="gt_row gt_left gt_striped" style="color: #999999;">gear</td>
<td headers="estimate" class="gt_row gt_right gt_striped" style="color: #999999;">0.0000</td>
<td headers="abs_estimate" class="gt_row gt_right gt_striped" style="color: #999999;">0.0000</td></tr>
    <tr><td headers="term" class="gt_row gt_left" style="color: #999999;">carb</td>
<td headers="estimate" class="gt_row gt_right" style="color: #999999;">0.0000</td>
<td headers="abs_estimate" class="gt_row gt_right" style="color: #999999;">0.0000</td></tr>
  </tbody>
  <tfoot>
    <tr class="gt_sourcenotes">
      <td class="gt_sourcenote" colspan="3">tidylearn | lasso (regression) | mpg ~ . | n = 32</td>
    </tr>
  </tfoot>
</table>
</div>

<p>Both plots are ggplot2 objects — <code class="language-plaintext highlighter-rouge">theme_minimal()</code>, consistent
aesthetics, and directly passable to <code class="language-plaintext highlighter-rouge">ggplotly()</code> or <code class="language-plaintext highlighter-rouge">ggsave()</code>. The
table is a <code class="language-plaintext highlighter-rouge">gt</code> object with the same consistent styling.</p>

<h3 id="without-tidylearn-2">Without tidylearn</h3>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">library</span><span class="p">(</span><span class="n">glmnet</span><span class="p">)</span><span class="w">

</span><span class="c1"># Prepare model matrix (glmnet doesn't accept formulas)</span><span class="w">
</span><span class="n">x</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">model.matrix</span><span class="p">(</span><span class="n">mpg</span><span class="w"> </span><span class="o">~</span><span class="w"> </span><span class="n">.</span><span class="p">,</span><span class="w"> </span><span class="n">data</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">mtcars</span><span class="p">)[,</span><span class="w"> </span><span class="m">-1</span><span class="p">]</span><span class="w">
</span><span class="n">y</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">mtcars</span><span class="o">$</span><span class="n">mpg</span><span class="w">

</span><span class="c1"># Fit with cross-validation</span><span class="w">
</span><span class="n">cv_fit</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">cv.glmnet</span><span class="p">(</span><span class="n">x</span><span class="p">,</span><span class="w"> </span><span class="n">y</span><span class="p">,</span><span class="w"> </span><span class="n">alpha</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="m">1</span><span class="p">)</span><span class="w">
</span></code></pre></div></div>

<p>The default coefficient path plot:</p>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">plot</span><span class="p">(</span><span class="n">cv_fit</span><span class="o">$</span><span class="n">glmnet.fit</span><span class="p">,</span><span class="w"> </span><span class="n">xvar</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"lambda"</span><span class="p">,</span><span class="w"> </span><span class="n">label</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="kc">TRUE</span><span class="p">)</span><span class="w">
</span></code></pre></div></div>

<p><img src="https://blog.sheetsolved.com/assets/tidylearn/base-lasso-path-1.png" alt="" /></p>

<p>The default cross-validation plot:</p>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">plot</span><span class="p">(</span><span class="n">cv_fit</span><span class="p">)</span><span class="w">
</span></code></pre></div></div>

<p><img src="https://blog.sheetsolved.com/assets/tidylearn/base-lasso-cv-1.png" alt="" /></p>

<p>These are base R graphics — functional, but they can’t be themed,
faceted, combined with other ggplot2 panels, or converted to interactive
plotly charts. Building ggplot2 equivalents from the <code class="language-plaintext highlighter-rouge">glmnet</code> object
requires extracting the coefficient matrix across all lambda values and
pivoting it to long format:</p>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># Extract coefficient matrix</span><span class="w">
</span><span class="n">coef_matrix</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">as.matrix</span><span class="p">(</span><span class="n">cv_fit</span><span class="o">$</span><span class="n">glmnet.fit</span><span class="o">$</span><span class="n">beta</span><span class="p">)</span><span class="w">
</span><span class="n">lambda_vals</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">cv_fit</span><span class="o">$</span><span class="n">glmnet.fit</span><span class="o">$</span><span class="n">lambda</span><span class="w">

</span><span class="c1"># Reshape to long format for ggplot</span><span class="w">
</span><span class="n">coef_df</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">as.data.frame</span><span class="p">(</span><span class="n">t</span><span class="p">(</span><span class="n">coef_matrix</span><span class="p">))</span><span class="w">
</span><span class="n">coef_df</span><span class="o">$</span><span class="n">lambda</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">lambda_vals</span><span class="w">
</span><span class="n">coef_long</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">tidyr</span><span class="o">::</span><span class="n">pivot_longer</span><span class="p">(</span><span class="n">coef_df</span><span class="p">,</span><span class="w"> </span><span class="o">-</span><span class="n">lambda</span><span class="p">,</span><span class="w">
                                  </span><span class="n">names_to</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"variable"</span><span class="p">,</span><span class="w">
                                  </span><span class="n">values_to</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"coefficient"</span><span class="p">)</span><span class="w">

</span><span class="n">ggplot</span><span class="p">(</span><span class="n">coef_long</span><span class="p">,</span><span class="w"> </span><span class="n">aes</span><span class="p">(</span><span class="n">x</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="nf">log</span><span class="p">(</span><span class="n">lambda</span><span class="p">),</span><span class="w"> </span><span class="n">y</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">coefficient</span><span class="p">,</span><span class="w"> </span><span class="n">colour</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">variable</span><span class="p">))</span><span class="w"> </span><span class="o">+</span><span class="w">
  </span><span class="n">geom_line</span><span class="p">()</span><span class="w"> </span><span class="o">+</span><span class="w">
  </span><span class="n">labs</span><span class="p">(</span><span class="n">title</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"Lasso Coefficient Path"</span><span class="p">,</span><span class="w"> </span><span class="n">x</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"log(lambda)"</span><span class="p">,</span><span class="w">
       </span><span class="n">y</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"Coefficient"</span><span class="p">,</span><span class="w"> </span><span class="n">colour</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"Variable"</span><span class="p">)</span><span class="w"> </span><span class="o">+</span><span class="w">
  </span><span class="n">theme_minimal</span><span class="p">()</span><span class="w">
</span></code></pre></div></div>

<p><img src="https://blog.sheetsolved.com/assets/tidylearn/base-lasso-gg-1.png" alt="" /></p>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># Extract coefficients at lambda.1se — returns a sparse matrix</span><span class="w">
</span><span class="n">coefs</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">as.matrix</span><span class="p">(</span><span class="n">coef</span><span class="p">(</span><span class="n">cv_fit</span><span class="p">,</span><span class="w"> </span><span class="n">s</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"lambda.1se"</span><span class="p">))</span><span class="w">
</span><span class="n">coef_tbl</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">data.frame</span><span class="p">(</span><span class="w">
  </span><span class="n">term</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">rownames</span><span class="p">(</span><span class="n">coefs</span><span class="p">),</span><span class="w">
  </span><span class="n">estimate</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">as.vector</span><span class="p">(</span><span class="n">coefs</span><span class="p">)</span><span class="w">
</span><span class="p">)</span><span class="w">
</span><span class="n">coef_tbl</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">coef_tbl</span><span class="p">[</span><span class="n">order</span><span class="p">(</span><span class="o">-</span><span class="nf">abs</span><span class="p">(</span><span class="n">coef_tbl</span><span class="o">$</span><span class="n">estimate</span><span class="p">)),</span><span class="w"> </span><span class="p">]</span><span class="w">

</span><span class="n">knitr</span><span class="o">::</span><span class="n">kable</span><span class="p">(</span><span class="n">coef_tbl</span><span class="p">,</span><span class="w"> </span><span class="n">digits</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="m">4</span><span class="p">,</span><span class="w"> </span><span class="n">row.names</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="kc">FALSE</span><span class="p">,</span><span class="w">
             </span><span class="n">caption</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"Lasso Coefficients at lambda.1se (manual)"</span><span class="p">)</span><span class="w">
</span></code></pre></div></div>

<table>
  <thead>
    <tr>
      <th style="text-align: left">term</th>
      <th style="text-align: right">estimate</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left">(Intercept)</td>
      <td style="text-align: right">33.9405</td>
    </tr>
    <tr>
      <td style="text-align: left">wt</td>
      <td style="text-align: right">-2.3659</td>
    </tr>
    <tr>
      <td style="text-align: left">cyl</td>
      <td style="text-align: right">-0.8430</td>
    </tr>
    <tr>
      <td style="text-align: left">hp</td>
      <td style="text-align: right">-0.0070</td>
    </tr>
    <tr>
      <td style="text-align: left">disp</td>
      <td style="text-align: right">0.0000</td>
    </tr>
    <tr>
      <td style="text-align: left">drat</td>
      <td style="text-align: right">0.0000</td>
    </tr>
    <tr>
      <td style="text-align: left">qsec</td>
      <td style="text-align: right">0.0000</td>
    </tr>
    <tr>
      <td style="text-align: left">vs</td>
      <td style="text-align: right">0.0000</td>
    </tr>
    <tr>
      <td style="text-align: left">am</td>
      <td style="text-align: right">0.0000</td>
    </tr>
    <tr>
      <td style="text-align: left">gear</td>
      <td style="text-align: right">0.0000</td>
    </tr>
    <tr>
      <td style="text-align: left">carb</td>
      <td style="text-align: right">0.0000</td>
    </tr>
  </tbody>
</table>

<p>Lasso Coefficients at lambda.1se (manual)</p>

<p>The manual approach works, but it’s the kind of reshaping code that’s
easy to get subtly wrong and tedious to repeat.
<code class="language-plaintext highlighter-rouge">tl_plot_regularization_path()</code> handles extraction, pivoting, labelling,
and theming in one call. And the <code class="language-plaintext highlighter-rouge">kable()</code> coefficient table is plain —
<code class="language-plaintext highlighter-rouge">tl_table_coefficients()</code> adds sorting, zero greying, and the selected
lambda value in the subtitle.</p>

<hr />

<h2 id="4-multi-model-classification-comparison">4. Multi-Model Classification Comparison</h2>

<p>Comparing models across packages is where consistent output structure
matters most. Each package has its own prediction interface, metric
accessors, and plot conventions.</p>

<h3 id="with-tidylearn-3">With tidylearn</h3>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">split</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">tl_split</span><span class="p">(</span><span class="n">iris</span><span class="p">,</span><span class="w"> </span><span class="n">prop</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="m">0.7</span><span class="p">,</span><span class="w"> </span><span class="n">stratify</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"Species"</span><span class="p">,</span><span class="w"> </span><span class="n">seed</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="m">42</span><span class="p">)</span><span class="w">

</span><span class="c1"># Fit three models — same interface for each</span><span class="w">
</span><span class="n">m_forest</span><span class="w">   </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">tl_model</span><span class="p">(</span><span class="n">split</span><span class="o">$</span><span class="n">train</span><span class="p">,</span><span class="w"> </span><span class="n">Species</span><span class="w"> </span><span class="o">~</span><span class="w"> </span><span class="n">.</span><span class="p">,</span><span class="w"> </span><span class="n">method</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"forest"</span><span class="p">)</span><span class="w">
</span><span class="n">m_tree</span><span class="w">     </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">tl_model</span><span class="p">(</span><span class="n">split</span><span class="o">$</span><span class="n">train</span><span class="p">,</span><span class="w"> </span><span class="n">Species</span><span class="w"> </span><span class="o">~</span><span class="w"> </span><span class="n">.</span><span class="p">,</span><span class="w"> </span><span class="n">method</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"tree"</span><span class="p">)</span><span class="w">
</span><span class="n">m_xgboost</span><span class="w">  </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">tl_model</span><span class="p">(</span><span class="n">split</span><span class="o">$</span><span class="n">train</span><span class="p">,</span><span class="w"> </span><span class="n">Species</span><span class="w"> </span><span class="o">~</span><span class="w"> </span><span class="n">.</span><span class="p">,</span><span class="w"> </span><span class="n">method</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"xgboost"</span><span class="p">)</span><span class="w">
</span></code></pre></div></div>

<p>A formatted comparison table — one call:</p>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">tl_table_comparison</span><span class="p">(</span><span class="w">
  </span><span class="n">m_forest</span><span class="p">,</span><span class="w"> </span><span class="n">m_tree</span><span class="p">,</span><span class="w"> </span><span class="n">m_xgboost</span><span class="p">,</span><span class="w">
  </span><span class="n">new_data</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">split</span><span class="o">$</span><span class="n">test</span><span class="p">,</span><span class="w">
  </span><span class="n">names</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="nf">c</span><span class="p">(</span><span class="s2">"Random Forest"</span><span class="p">,</span><span class="w"> </span><span class="s2">"Decision Tree"</span><span class="p">,</span><span class="w"> </span><span class="s2">"XGBoost"</span><span class="p">)</span><span class="w">
</span><span class="p">)</span><span class="w">
</span></code></pre></div></div>

<div id="xtzejhxvix" class="gt-post-table" style="padding-left:0px;padding-right:0px;padding-top:10px;padding-bottom:10px;overflow-x:auto;overflow-y:auto;width:auto;height:auto;">

<table class="gt_table" data-quarto-disable-processing="false" data-quarto-bootstrap="false">
  <thead>
    <tr class="gt_heading">
      <td colspan="4" class="gt_heading gt_title gt_font_normal" style="">Model Comparison</td>
    </tr>
    <tr class="gt_heading">
      <td colspan="4" class="gt_heading gt_subtitle gt_font_normal gt_bottom_border" style="">3 models compared</td>
    </tr>
    <tr class="gt_col_headings">
      <th class="gt_col_heading gt_columns_bottom_border gt_left" rowspan="1" colspan="1" style="color: #FFFFFF;" scope="col" id="metric">Metric</th>
      <th class="gt_col_heading gt_columns_bottom_border gt_right" rowspan="1" colspan="1" style="color: #FFFFFF;" scope="col" id="Random-Forest">Random Forest</th>
      <th class="gt_col_heading gt_columns_bottom_border gt_right" rowspan="1" colspan="1" style="color: #FFFFFF;" scope="col" id="Decision-Tree">Decision Tree</th>
      <th class="gt_col_heading gt_columns_bottom_border gt_right" rowspan="1" colspan="1" style="color: #FFFFFF;" scope="col" id="XGBoost">XGBoost</th>
    </tr>
  </thead>
  <tbody class="gt_table_body">
    <tr><td headers="metric" class="gt_row gt_left">Accuracy</td>
<td headers="Random Forest" class="gt_row gt_right">0.9556</td>
<td headers="Decision Tree" class="gt_row gt_right">0.8889</td>
<td headers="XGBoost" class="gt_row gt_right">0.9111</td></tr>
  </tbody>
  <tfoot>
    <tr class="gt_sourcenotes">
      <td class="gt_sourcenote" colspan="4">tidylearn | n = 45</td>
    </tr>
  </tfoot>
</table>
</div>

<p>And a confusion matrix for any model:</p>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">tl_table_confusion</span><span class="p">(</span><span class="n">m_forest</span><span class="p">,</span><span class="w"> </span><span class="n">new_data</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">split</span><span class="o">$</span><span class="n">test</span><span class="p">)</span><span class="w">
</span></code></pre></div></div>

<div id="zcikultaqv" class="gt-post-table" style="padding-left:0px;padding-right:0px;padding-top:10px;padding-bottom:10px;overflow-x:auto;overflow-y:auto;width:auto;height:auto;">

<table class="gt_table" data-quarto-disable-processing="false" data-quarto-bootstrap="false">
  <thead>
    <tr class="gt_heading">
      <td colspan="4" class="gt_heading gt_title gt_font_normal gt_bottom_border" style="">Confusion Matrix</td>
    </tr>
    
    <tr class="gt_col_headings gt_spanner_row">
      <th class="gt_col_heading gt_columns_bottom_border gt_left" rowspan="2" colspan="1" scope="col" id="a::stub">Actual</th>
      <th class="gt_center gt_columns_top_border gt_column_spanner_outer" rowspan="1" colspan="3" scope="colgroup" id="Predicted">
        <div class="gt_column_spanner">Predicted</div>
      </th>
    </tr>
    <tr class="gt_col_headings">
      <th class="gt_col_heading gt_columns_bottom_border gt_right" rowspan="1" colspan="1" style="color: #FFFFFF;" scope="col" id="setosa">setosa</th>
      <th class="gt_col_heading gt_columns_bottom_border gt_right" rowspan="1" colspan="1" style="color: #FFFFFF;" scope="col" id="versicolor">versicolor</th>
      <th class="gt_col_heading gt_columns_bottom_border gt_right" rowspan="1" colspan="1" style="color: #FFFFFF;" scope="col" id="virginica">virginica</th>
    </tr>
  </thead>
  <tbody class="gt_table_body">
    <tr><th id="stub_1_1" scope="row" class="gt_row gt_left gt_stub">setosa</th>
<td headers="stub_1_1 setosa" class="gt_row gt_right" style="background-color: #D4EDDA;">15</td>
<td headers="stub_1_1 versicolor" class="gt_row gt_right">0</td>
<td headers="stub_1_1 virginica" class="gt_row gt_right">0</td></tr>
    <tr><th id="stub_1_2" scope="row" class="gt_row gt_left gt_stub">versicolor</th>
<td headers="stub_1_2 setosa" class="gt_row gt_right gt_striped">0</td>
<td headers="stub_1_2 versicolor" class="gt_row gt_right gt_striped" style="background-color: #D4EDDA;">15</td>
<td headers="stub_1_2 virginica" class="gt_row gt_right gt_striped">0</td></tr>
    <tr><th id="stub_1_3" scope="row" class="gt_row gt_left gt_stub">virginica</th>
<td headers="stub_1_3 setosa" class="gt_row gt_right">0</td>
<td headers="stub_1_3 versicolor" class="gt_row gt_right">2</td>
<td headers="stub_1_3 virginica" class="gt_row gt_right" style="background-color: #D4EDDA;">13</td></tr>
  </tbody>
  <tfoot>
    <tr class="gt_sourcenotes">
      <td class="gt_sourcenote" colspan="4">tidylearn | forest (classification) | Species ~ . | n = 105</td>
    </tr>
  </tfoot>
</table>
</div>

<p>Adding a fourth model is just another argument in
<code class="language-plaintext highlighter-rouge">tl_table_comparison()</code> — the table code stays unchanged.</p>

<h3 id="without-tidylearn-3">Without tidylearn</h3>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">library</span><span class="p">(</span><span class="n">randomForest</span><span class="p">)</span><span class="w">
</span><span class="n">library</span><span class="p">(</span><span class="n">xgboost</span><span class="p">)</span><span class="w">

</span><span class="n">set.seed</span><span class="p">(</span><span class="m">42</span><span class="p">)</span><span class="w">
</span><span class="n">train_idx</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">unlist</span><span class="p">(</span><span class="n">lapply</span><span class="p">(</span><span class="w">
  </span><span class="n">split</span><span class="p">(</span><span class="nf">seq_len</span><span class="p">(</span><span class="n">nrow</span><span class="p">(</span><span class="n">iris</span><span class="p">)),</span><span class="w"> </span><span class="n">iris</span><span class="o">$</span><span class="n">Species</span><span class="p">),</span><span class="w">
  </span><span class="k">function</span><span class="p">(</span><span class="n">i</span><span class="p">)</span><span class="w"> </span><span class="n">sample</span><span class="p">(</span><span class="n">i</span><span class="p">,</span><span class="w"> </span><span class="n">size</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="nf">floor</span><span class="p">(</span><span class="m">0.7</span><span class="w"> </span><span class="o">*</span><span class="w"> </span><span class="nf">length</span><span class="p">(</span><span class="n">i</span><span class="p">)))</span><span class="w">
</span><span class="p">))</span><span class="w">
</span><span class="n">train_data</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">iris</span><span class="p">[</span><span class="n">train_idx</span><span class="p">,</span><span class="w"> </span><span class="p">]</span><span class="w">
</span><span class="n">test_data</span><span class="w">  </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">iris</span><span class="p">[</span><span class="o">-</span><span class="n">train_idx</span><span class="p">,</span><span class="w"> </span><span class="p">]</span><span class="w">

</span><span class="c1"># Two of the three take a formula and a data frame</span><span class="w">
</span><span class="n">fit_rf</span><span class="w">   </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">randomForest</span><span class="p">(</span><span class="n">Species</span><span class="w"> </span><span class="o">~</span><span class="w"> </span><span class="n">.</span><span class="p">,</span><span class="w"> </span><span class="n">data</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">train_data</span><span class="p">)</span><span class="w">
</span><span class="n">fit_tree</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">rpart</span><span class="o">::</span><span class="n">rpart</span><span class="p">(</span><span class="n">Species</span><span class="w"> </span><span class="o">~</span><span class="w"> </span><span class="n">.</span><span class="p">,</span><span class="w"> </span><span class="n">data</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">train_data</span><span class="p">,</span><span class="w"> </span><span class="n">method</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"class"</span><span class="p">)</span><span class="w">

</span><span class="c1"># xgboost takes neither: a numeric matrix, integer-encoded labels, and</span><span class="w">
</span><span class="c1"># xgb.train() rather than xgboost(), which refuses multiclass outright</span><span class="w">
</span><span class="n">dtrain</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">xgb.DMatrix</span><span class="p">(</span><span class="w">
  </span><span class="n">data</span><span class="w">  </span><span class="o">=</span><span class="w"> </span><span class="n">as.matrix</span><span class="p">(</span><span class="n">train_data</span><span class="p">[,</span><span class="w"> </span><span class="m">1</span><span class="o">:</span><span class="m">4</span><span class="p">]),</span><span class="w">
  </span><span class="n">label</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="nf">as.integer</span><span class="p">(</span><span class="n">train_data</span><span class="o">$</span><span class="n">Species</span><span class="p">)</span><span class="w"> </span><span class="o">-</span><span class="w"> </span><span class="m">1</span><span class="w">
</span><span class="p">)</span><span class="w">
</span><span class="n">fit_xgb</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">xgb.train</span><span class="p">(</span><span class="w">
  </span><span class="n">params</span><span class="w">  </span><span class="o">=</span><span class="w"> </span><span class="nf">list</span><span class="p">(</span><span class="n">objective</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"multi:softmax"</span><span class="p">,</span><span class="w"> </span><span class="n">num_class</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="m">3</span><span class="p">),</span><span class="w">
  </span><span class="n">data</span><span class="w">    </span><span class="o">=</span><span class="w"> </span><span class="n">dtrain</span><span class="p">,</span><span class="w">
  </span><span class="n">nrounds</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="m">20</span><span class="p">,</span><span class="w">
  </span><span class="n">verbose</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="m">0</span><span class="w">
</span><span class="p">)</span><span class="w">

</span><span class="c1"># Each returns predictions in a different format</span><span class="w">
</span><span class="n">pred_rf</span><span class="w">   </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">predict</span><span class="p">(</span><span class="n">fit_rf</span><span class="p">,</span><span class="w"> </span><span class="n">newdata</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">test_data</span><span class="p">)</span><span class="w">
</span><span class="n">pred_tree</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">predict</span><span class="p">(</span><span class="n">fit_tree</span><span class="p">,</span><span class="w"> </span><span class="n">newdata</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">test_data</span><span class="p">,</span><span class="w"> </span><span class="n">type</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"class"</span><span class="p">)</span><span class="w">
</span><span class="n">pred_xgb</span><span class="w">  </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">predict</span><span class="p">(</span><span class="n">fit_xgb</span><span class="p">,</span><span class="w"> </span><span class="n">xgb.DMatrix</span><span class="p">(</span><span class="n">as.matrix</span><span class="p">(</span><span class="n">test_data</span><span class="p">[,</span><span class="w"> </span><span class="m">1</span><span class="o">:</span><span class="m">4</span><span class="p">])))</span><span class="w">

</span><span class="c1"># And the xgboost predictions are zero-based integers, not factor levels</span><span class="w">
</span><span class="n">acc_rf</span><span class="w">   </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">mean</span><span class="p">(</span><span class="n">pred_rf</span><span class="w"> </span><span class="o">==</span><span class="w"> </span><span class="n">test_data</span><span class="o">$</span><span class="n">Species</span><span class="p">)</span><span class="w">
</span><span class="n">acc_tree</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">mean</span><span class="p">(</span><span class="n">pred_tree</span><span class="w"> </span><span class="o">==</span><span class="w"> </span><span class="n">test_data</span><span class="o">$</span><span class="n">Species</span><span class="p">)</span><span class="w">
</span><span class="n">acc_xgb</span><span class="w">  </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">mean</span><span class="p">(</span><span class="n">levels</span><span class="p">(</span><span class="n">iris</span><span class="o">$</span><span class="n">Species</span><span class="p">)[</span><span class="n">pred_xgb</span><span class="w"> </span><span class="o">+</span><span class="w"> </span><span class="m">1</span><span class="p">]</span><span class="w"> </span><span class="o">==</span><span class="w"> </span><span class="n">test_data</span><span class="o">$</span><span class="n">Species</span><span class="p">)</span><span class="w">

</span><span class="n">comparison_base</span><span class="w"> </span><span class="o">&lt;-</span><span class="w"> </span><span class="n">data.frame</span><span class="p">(</span><span class="w">
  </span><span class="n">model</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="nf">c</span><span class="p">(</span><span class="s2">"Random Forest"</span><span class="p">,</span><span class="w"> </span><span class="s2">"Decision Tree"</span><span class="p">,</span><span class="w"> </span><span class="s2">"XGBoost"</span><span class="p">),</span><span class="w">
  </span><span class="n">accuracy</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="nf">c</span><span class="p">(</span><span class="n">acc_rf</span><span class="p">,</span><span class="w"> </span><span class="n">acc_tree</span><span class="p">,</span><span class="w"> </span><span class="n">acc_xgb</span><span class="p">)</span><span class="w">
</span><span class="p">)</span><span class="w">

</span><span class="n">knitr</span><span class="o">::</span><span class="n">kable</span><span class="p">(</span><span class="n">comparison_base</span><span class="p">,</span><span class="w"> </span><span class="n">digits</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="m">3</span><span class="p">,</span><span class="w">
             </span><span class="n">caption</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"Model Comparison (manual)"</span><span class="p">)</span><span class="w">
</span></code></pre></div></div>

<table>
  <thead>
    <tr>
      <th style="text-align: left">model</th>
      <th style="text-align: right">accuracy</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left">Random Forest</td>
      <td style="text-align: right">0.933</td>
    </tr>
    <tr>
      <td style="text-align: left">Decision Tree</td>
      <td style="text-align: right">0.889</td>
    </tr>
    <tr>
      <td style="text-align: left">XGBoost</td>
      <td style="text-align: right">0.911</td>
    </tr>
  </tbody>
</table>

<p>Model Comparison (manual)</p>

<p>Three packages, three interfaces. <code class="language-plaintext highlighter-rouge">randomForest</code> and <code class="language-plaintext highlighter-rouge">rpart</code> take a
formula and a data frame; <code class="language-plaintext highlighter-rouge">xgboost</code> takes a numeric matrix wrapped in a
<code class="language-plaintext highlighter-rouge">DMatrix</code>, integer-encoded labels, and <code class="language-plaintext highlighter-rouge">xgb.train()</code> rather than
<code class="language-plaintext highlighter-rouge">xgboost()</code>, which refuses multiclass objectives outright. The
predictions come back as factor levels, factor levels, and zero-based
integers respectively, so each accuracy has to be computed its own way.
The <code class="language-plaintext highlighter-rouge">kable()</code> output is plain and limited to one metric.</p>

<p><code class="language-plaintext highlighter-rouge">tl_table_comparison()</code> takes the fitted models and produces a styled,
multi-metric table without any of that reshaping.</p>

<hr />

<h2 id="5-interactive-reporting-with-plotly">5. Interactive Reporting with plotly</h2>

<p>Because tidylearn’s plot functions return standard ggplot2 objects,
converting any visualisation to an interactive plotly chart is a
one-liner:</p>

<div class="language-r highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">library</span><span class="p">(</span><span class="n">plotly</span><span class="p">)</span><span class="w">

</span><span class="c1"># tidylearn's plot returns a ggplot2 object — pass it straight to ggplotly</span><span class="w">
</span><span class="n">ggplotly</span><span class="p">(</span><span class="n">tidy_pca_biplot</span><span class="p">(</span><span class="n">pca</span><span class="p">,</span><span class="w"> </span><span class="n">label_obs</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="kc">TRUE</span><span class="p">))</span><span class="w">
</span><span class="n">ggplotly</span><span class="p">(</span><span class="n">tl_plot_regularization_path</span><span class="p">(</span><span class="n">lasso</span><span class="p">))</span><span class="w">
</span><span class="n">ggplotly</span><span class="p">(</span><span class="n">plot</span><span class="p">(</span><span class="n">m_forest</span><span class="p">,</span><span class="w"> </span><span class="n">type</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="s2">"confusion"</span><span class="p">))</span><span class="w">
</span></code></pre></div></div>

<p><code class="language-plaintext highlighter-rouge">ggplotly()</code> picks up axis labels, themes, and tooltip data
automatically. Compare this to the base R plots from <code class="language-plaintext highlighter-rouge">biplot()</code>,
<code class="language-plaintext highlighter-rouge">plot.glmnet()</code>, or <code class="language-plaintext highlighter-rouge">plot.hclust()</code> — none of which can be converted to
plotly without rebuilding them from scratch.</p>

<hr />

<h2 id="what-the-uniformity-buys">What the Uniformity Buys</h2>

<p><strong>Consistent, polished output by default.</strong> Every model — whether it’s a
PCA biplot, a lasso coefficient path, or a confusion matrix — returns
ggplot2 plots and <code class="language-plaintext highlighter-rouge">gt</code> tables with a consistent visual language. You
don’t need to learn each package’s idiosyncratic output format or build
custom formatting code to get report-quality visuals and tables.</p>

<p><strong>Reproducibility through uniformity.</strong> When your reporting pipeline
works the same way for every model type, your analysis becomes genuinely
reproducible. Swap <code class="language-plaintext highlighter-rouge">method = "forest"</code> for <code class="language-plaintext highlighter-rouge">method = "xgboost"</code> and
rerun — the same <code class="language-plaintext highlighter-rouge">tl_table()</code> calls, the same <code class="language-plaintext highlighter-rouge">plot()</code> calls, the same
comparison logic all work without modification. That means you can
iterate on model selection without touching your reporting layer, and
anyone reading your code can follow the same pattern across different
analyses.</p>

<p>The best analysis code is code that gets out of your way and lets you
focus on the results. That’s what tidylearn is for.</p>]]></content><author><name>Cesaire Tobias</name></author><category term="R" /><category term="r" /><category term="machine-learning" /><category term="tidylearn" /><summary type="html"><![CDATA[Every R modelling package returns results in a different shape, so reporting code breaks whenever the model changes. What one tidy interface across twenty algorithms buys you.]]></summary></entry><entry><title type="html">Partly Cloudy: Forecasting Your Needs in a Fragmented Cloud</title><link href="https://blog.sheetsolved.com/partly-cloudy.html" rel="alternate" type="text/html" title="Partly Cloudy: Forecasting Your Needs in a Fragmented Cloud" /><published>2026-06-02T00:00:00+00:00</published><updated>2026-06-02T00:00:00+00:00</updated><id>https://blog.sheetsolved.com/partly-cloudy</id><content type="html" xml:base="https://blog.sheetsolved.com/partly-cloudy.html"><![CDATA[<p>You can stand up a serious production stack today without ever opening a hyperscaler console. GPU compute from a neocloud, object storage from Cloudflare, ephemeral functions from Modal, deployment from Vercel — each piece a few lines of config, each one live in minutes. Ten years ago that stack was either science fiction or a maintenance burden no small team would take on. Now it’s an ordinary afternoon’s work.</p>

<p>That shift is easy to miss under the louder noise about AI, but it matters at least as much. The cloud has come apart into components, and the consequence reaches past the technical. It changes what the central question even is. For a long time the question was <em>can we run this at all?</em> Now it’s <em>which of dozens of viable options do we pick, and what does it cost us to keep them talking to each other?</em></p>

<h2 id="commitment-became-composition">Commitment became composition</h2>

<p>There was a time when reaching for the cloud was a weighty, deliberate act. The cloud was for big things — real storage, real compute, the workloads you genuinely couldn’t run on your own hardware. Getting there meant learning networking, identity, regions, billing, a stack of concepts you had to absorb before a single line of your own code would run.</p>

<p>That overhead did something subtle. It pushed everyone toward standardising on one provider. If you were going to pay the setup cost once, you paid it on AWS, you learned AWS, and you did <em>everything</em> on AWS. The friction was itself a centralising force. Nobody shopped around, because shopping around meant paying the tax again.</p>

<p>The friction is mostly gone now. Provisioning a GPU, a bucket, or a function is close to free in both money and effort. And when the cost of trying something drops to nearly zero, the cloud stops being a destination you commit to and becomes a set of parts you compose. You no longer adopt a specialist because there’s no other way to do the job. You adopt it because, for that one slice of the problem, it is faster, cheaper, or simply more pleasant than the generalist. The cloud went from something you needed to something you reach for.</p>

<h2 id="why-it-came-apart">Why it came apart</h2>

<p>The single broad platform has been unbundled into a field of specialists, each one taking a single layer and beating the generalist inside it.</p>

<p>Neoclouds like Nebius, CoreWeave, and Lambda took GPU compute, born out of a stretch where the hyperscalers simply couldn’t supply accelerators fast enough or at a sane price. Cloudflare’s R2 took object storage, leading with no egress fees — a direct strike at what was, on the big platforms, less a cost than a way to make leaving expensive. Modal took ephemeral compute, Vercel took deployment, and there is a new name worth evaluating most months.</p>

<p>Each of these succeeds by being deep instead of broad. They don’t try to be all things to everyone. They try to be the best available version of one thing, and they tend to arrive bottom-up — a developer adopts one to solve a problem in front of them — rather than top-down through an architecture review.</p>

<h2 id="the-bill-you-dont-see-at-signup">The bill you don’t see at signup</h2>

<p>A perfect tool for every layer is wonderful right up to the point where you have to assemble them.</p>

<p>Five specialists means five identity models, five billing relationships, five security postures to reason about, data moving <em>between</em> providers, and observability stitched across boundaries that don’t naturally line up. The specialist saves money on the line item and bills you back, quietly, in integration work. This is the moment the generalist’s perfectly serviceable managed database starts to look appealing — not because it’s better, but because it is already there and already wired into everything else.</p>

<p>There is a second cost, subtler than integration and rarely mentioned: the cost of <em>knowing</em>. When a credible new platform appears most months, how do you know when something genuinely better has arrived? How much time goes into evaluating it, and how much churn can a team absorb chasing improvements that are real but marginal? The freedom to pick the best tool becomes a standing obligation to keep checking whether you still have it. That is a real and continuous load, and it didn’t exist when the friction made the decision once and left it settled.</p>

<h2 id="when-the-platform-still-wins">When the platform still wins</h2>

<p>None of this retires the generalist. It relocates its value.</p>

<p>The hyperscaler used to be valuable because it was the only practical option. That is over, but what remains is genuinely worth paying for, to the right buyer. Integration is the actual product: one identity model, one invoice, services built to wire into each other, so that if your binding constraint is engineering time rather than the cloud bill, the coherence is worth more than a cheaper component. Procurement and compliance matter too — committed-spend agreements, a single vendor to hold accountable, broad certification — which is unglamorous and exactly why large organisations still default there. And breadth is a form of insurance: you don’t re-evaluate vendors every time your needs shift, and good-enough-and-present beats excellent-but-elsewhere more often than engineers like to admit.</p>

<h2 id="the-view-from-here">The view from here</h2>

<p>Most of that argument has been written as though the choice were borderless. From where we sit, it isn’t. Geography reasserts itself the moment you take residency, latency, and regulation seriously — and in South Africa all three arrive at once.</p>

<p>The hyperscalers have spent real money locally. AWS <a href="https://press.aboutamazon.com/2020/4/aws-launches-region-in-south-africa">opened its Cape Town region in 2020</a>, Microsoft runs regions in both Johannesburg and Cape Town, and Google <a href="https://cloud.google.com/blog/products/infrastructure/heita-south-africa-new-cloud-region">opened Johannesburg in 2024</a>. That means in-country data residency and low latency from a generalist you can also buy everything else from. Most of the new specialists offer nothing of the sort: the GPU neoclouds and platforms like Modal run from data centres abroad, with no African region at all. A few are exceptions — Cloudflare has <a href="https://www.cloudflare.com/network/">edge presence in Johannesburg, Cape Town, and Durban</a> — but the typical specialist sits across an ocean. Using one from here means accepting the latency and, more consequentially, sending your data offshore.</p>

<p>That last part is where the trade-off stops being academic. POPIA doesn’t demand blanket localisation, but <a href="https://popia.co.za/section-72-transfers-of-personal-information-outside-republic/">section 72</a> restricts moving personal information out of the country without adequate protection in place, and the regulators who watch finance and insurance lean hard toward keeping customer data close to home. For a regulated client, a specialist with no African region turns a latency question into a compliance conversation, with penalties that run to millions of rands. Layer on the currency exposure of paying for a USD-billed service in rands, and the marginal performance edge that made the specialist attractive can quietly be eaten alive by everything attached to it.</p>

<p>So the broad-versus-deep call tilts here in a way it doesn’t in the US. A specialist’s depth has to clear a higher bar — enough advantage to justify the latency, the cross-border data question, and the currency risk all at once. Sometimes it does. Often, for the regulated and the latency-sensitive, the local region of a generalist wins on exactly the grounds the global commentary overlooks. Knowing which case you’re in isn’t a generic skill, and it doesn’t come off a pricing page.</p>

<h2 id="whats-temporary-and-what-isnt">What’s temporary and what isn’t</h2>

<p>Some of this fragmentation will compress, and some won’t, and it pays to tell them apart.</p>

<p>The neocloud boom is partly an arbitrage on a supply shortage. The clearest sign of how unusual that shortage is: the hyperscalers themselves have become customers of the specialists, renting capacity they can’t build out fast enough. When supply normalises and the platforms’ own silicon matures, some of that margin will close, and some of these names will be absorbed or undercut.</p>

<p>The egress story is the opposite. It started as a competitive pitch and in the EU it is becoming law. From January 2027 <a href="https://digital-strategy.ec.europa.eu/en/factpages/data-act-explained">the Data Act</a> bars cloud providers from charging switching fees at all — including the egress charges that made leaving expensive — dismantling through regulation a lock-in that competition had only dented. A moat that is being removed by statute is not coming back. The same goes for bottom-up adoption: once developers can route around procurement to pick their own tools, that habit doesn’t reverse. One of these forces is a passing distortion. The others are a change in the terrain.</p>

<h2 id="the-skill-moved">The skill moved</h2>

<p>The cloud stopped being a decision you make once and became one you make continuously.</p>

<p>The old skill was provisioning — could you stand the thing up at all? That has been commoditised, and AI tooling has finished the job. The skill that matters now is judgement: choosing well from an overwhelming menu, recognising when a marginal gain isn’t worth the integration cost, and holding the discipline to stop chasing the newest name once what you have is good enough. Knowing what to compose, what to consolidate, and when “already integrated” beats “technically better” is harder to teach and harder to hire for than knowing how to configure a VPC. Increasingly it is the work itself — less about standing systems up than about deciding what belongs where, and making the parts hold together once they’re chosen.</p>

<p>The barrier has shifted from <em>can you build it</em> to <em>can you choose well without drowning in the options</em>. That second skill is the one worth investing in now, and we’re not convinced it’s the easier of the two.</p>

<p>The cloud was made easier precisely so you wouldn’t need a specialist to use it — and the proliferation that came with that ease has made experienced guidance more valuable. What’s become scarce is judgement. When there was one obvious provider, advice was nearly redundant — you learned the platform and got on with it. With dozens of defensible options and a new one most months, the rare and useful thing is someone who has seen enough of these decisions to say which fragments are worth assembling, which are noise, and how to keep the whole thing coherent a year from now. Accessibility moved the need for expertise up the stack.</p>

<hr />

<p>June 2, 2026</p>]]></content><author><name>Cesaire Tobias</name></author><category term="cloud" /><category term="infrastructure" /><summary type="html"><![CDATA[The cloud has come apart into components. When a production stack needs no hyperscaler console, the question stops being whether you can run it and becomes which option to choose.]]></summary></entry><entry><title type="html">The Average Human Problem: Why AI “Sounds Like AI”</title><link href="https://blog.sheetsolved.com/ai-average-human-problem.html" rel="alternate" type="text/html" title="The Average Human Problem: Why AI “Sounds Like AI”" /><published>2026-05-04T00:00:00+00:00</published><updated>2026-05-04T00:00:00+00:00</updated><id>https://blog.sheetsolved.com/ai-average-human-problem</id><content type="html" xml:base="https://blog.sheetsolved.com/ai-average-human-problem.html"><![CDATA[<p>I keep seeing critiques of AI-generated writing where I think — but that’s how I would have written that. The em-dash. The tripled list. The “it’s not just X, it’s Y” cadence. The careful hedging. Even the use of bullet points! People point at these as if they’re machine fingerprints, and increasingly I want to ask: whose writing do you think the machine learned from in the first place?</p>

<p>The “AI tells” people detect are artefacts learned from us. But there’s something interesting in which of us — in which patterns get smoothed out, which get amplified, and who ends up being accused because of it.</p>

<h2 id="the-training-data-is-us">The Training Data Is Us</h2>

<p>Every pattern AI produces was learned from human text. Not human-adjacent or synthesised. You and me. The em-dash, the careful hedging, the structured paragraph with a thesis sentence and supporting beats — these weren’t invented by a model. They came from books, papers, blogs, essays, forum posts, readmes and repos.</p>

<p>When someone declares “this reads as AI,” they are pattern-matching against their own register and mistaking unfamiliarity for artificiality. The “AI tells” they cite aren’t tells of machinehood. They’re tells of careful prose — the kind written in academia, in formal correspondence, in essays by people taught to structure their thoughts. The kind written by people who learned English as a second language and overcorrected toward formality. The kind written by anyone who took a writing class seriously. Machines didn’t invent these patterns but we’re now seeing them being echoed at scale.</p>

<h2 id="average-humanness">Average Humanness</h2>

<p>AI doesn’t attempt to write like any single human. It writes like the centroid of millions of them. Every phrase is human-derived, but the combination trends toward median register — the idiosyncrasies, dead-ends, half-finished thoughts, and voice quirks that mark individual writers get smoothed out in averaging. Then RLHF — the human-feedback training that shapes how models respond — adds a second layer on top: a learned preference for structure, hedging, helpful framing and slightly formal tone.</p>

<p>So the “AI flavour” critics detect is not imagined but it isn’t fake humanness either. It’s averaged humanness. Which is a more interesting target than “AI sounds unnatural,” because it explains why the false-positive pattern lands where it does.</p>

<p>The people who write closest to that average, careful-prose baseline are the ones most likely to get caught in the dragnet. Their writing is closest to what AI was trained to produce.</p>

<h2 id="what-readers-are-actually-detecting">What Readers Are Actually Detecting</h2>

<p>If you write for a living, or think for a living, the loud version of this matters less than the discourse around it suggests. The inverse of the accusation is nearer the mark: the patterns flagged as machine-generated are the patterns of the most common kind of human writing. AI didn’t invent them — it learned them from us.</p>

<p>But by now we’ve all read AI-generated content, often without flagging it as such. What we actually care about, in the moment of reading, is whether the substance holds up and whether we enjoyed it — or, for more formal content, whether it lands clearly and without friction. When those are there, the question of provenance fades into the background. When they aren’t, the source — human or machine — doesn’t save it.</p>

<p>There’s a fair caveat. In creative or literary writing, sounding like the average is itself a problem. Voice is the point, and AI’s centroid-tendency is a real limitation there. But for the vast majority of functional writing, readability and substance are what we actually grade against.</p>

<p>So the scrutiny gap, in the end, is a highly subjective one. What readers are really trying to detect is probably evidence of thought. Whether the writer engaged with what they produced or just shipped a half-baked prompt. That’s the fair version of “this reads as AI.” The unfair version mistakes careful human prose for the same absence. The test we actually run when we’re reading is simpler than the one we invoke when we’re complaining: was this worth reading?</p>

<hr />

<p>May 4, 2026</p>]]></content><author><name>Cesaire Tobias</name></author><category term="ai" /><category term="writing" /><summary type="html"><![CDATA[The em-dash, the tripled list, the careful hedging: the tells people read as machine fingerprints were learned from us. On whose writing gets smoothed out, and who gets accused.]]></summary></entry></feed>