# Dr. Randal S. Olson - Full Content > AI Researcher & Builder. Co-Founder & CTO at Goodeye Labs, where we build AI products that point frontier models at the business outcomes that actually matter. Creator of TPOT (early AutoML tool) and 50+ peer-reviewed publications in AI/ML. The blog covers data visualization, data science, machine learning, and applied AI. Last generated: 2026-08-07 --- ## The people who use AI the most are also the most worried about it URL: https://www.randalolson.com/2026/06/25/ai-heaviest-users-most-worried/ Published: 2026-06-25 Categories: data visualization Tags: beautiful-charts-with-ai, artificial intelligence, public opinion, technology Pew's 2026 data shows the youngest U.S. adults use AI chatbots the most, yet they are also the most likely to expect AI to harm society. Part of Teaching an AI Agent to Make Beautiful Charts There is a comforting story we tell about new technology: the people who use it the most stop fearing it, because they learn what it can and cannot do. The latest numbers on AI break that story in half. The U.S. adults who use AI chatbots the most are also the ones most convinced AI will do harm. Pew Research Center asked more than 5,000 U.S. adults both halves of the question at once: do you use AI, and do you think it will help or hurt. Line the answers up by age and use and worry do not move together at all. Use of chatbots falls off a cliff as people get older. Worry barely budges. Each row is an age group. The blue dot is the share who have used an AI chatbot like ChatGPT, Gemini, or Copilot; the red dot is the share who think AI will do more harm than good to society over the next 20 years. The bar between them is the gap. Going down the rows, the blue dot slides far to the left while the red dot barely moves, so the bars shrink and, at the bottom, flip. The youngest adults are the heaviest users 66% of U.S. adults under 30 have used an AI chatbot, the highest of any age group, against just 23% of those 65 and older. Overall, about half of U.S. adults (49%) now use these tools, up from 33% in 2024. AI went mainstream fast, and it went mainstream youngest-first. The gap is not only about trying it once. Among under-50s, daily use clusters near the top: about 31% of under-30s and 34% of 30-to-49-year-olds reach for a chatbot every day. After 50 the habit thins out quickly. They are also the most worried 48% of under-30s say AI will harm society, the highest share of any age group, and only 14% expect it to help. The cohort that adopted AI fastest is the most pessimistic about where it leads. This is not really about whether the chatbot gives a good answer. People who lean on these tools every day mostly find them useful. The unease is bigger than the tool: 63% of all U.S. adults say AI is advancing too quickly, and 71% think it will make their personal information less secure. Heavy users see the upside and the downside up close, and the downside is the part that sticks. The heaviest users also have the most to lose Why would the people who use AI most be the most uneasy about it? Look at who is standing in front of the technology. Under-30s are not just the heaviest users; they are the most likely to think AI will hurt them personally, at 37%, against 28% of those 50 and older. That tracks the job market they are walking into. Unemployment among recent U.S. college graduates has climbed to nearly 6%, rising about twice as fast as for workers overall, and it bites hardest in the fields most exposed to automation, where new computer science graduates are around 7%. It is the same first rung recent graduates have been struggling to reach. And in the most AI-exposed jobs, employment for 22-to-25-year-olds has fallen while every older group held steady. The youngest workers are being asked to compete with a tool that does the starter tasks for free, so it adds up that they watch it most warily. The order flips: older adults barely use AI, but still worry By 65 and up, more U.S. adults expect AI to harm society (35%) than have ever used a chatbot (23%). That is the only row where the dots swap order. Use drops from 66% of the youngest group to 23% of the oldest, while worry only slips from 48% to 35%. Older adults' unease is more abstract. They are less likely to say AI will hurt them personally and more likely to worry about what it does to everyone else: scams, misinformation, the sense that something big is shifting without their say. You do not have to use a chatbot to be wary of one. How this chart was made This chart was built by an AI agent and graded against the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: Pew Research Center, "Americans and AI 2026", based on a survey of 5,119 U.S. adults conducted February 17-23, 2026. The usage and impact figures by age come from Pew's age breakdown. The cleaned dataset is available here. --- ## In AI-exposed jobs, only the youngest workers are losing ground URL: https://www.randalolson.com/2026/06/22/ai-jobs-hit-youngest-workers/ Published: 2026-06-22 Categories: data visualization Tags: beautiful-charts-with-ai, artificial intelligence, labor market, young workers, employment Since ChatGPT, employment for U.S. workers aged 22 to 25 in the most AI-exposed jobs has fallen about 12%, while every older age group held steady or grew. Part of Teaching an AI Agent to Make Beautiful Charts If AI were quietly thinning out jobs, you would expect the damage to spread across a workforce, not concentrate in a single corner of it. That is not what the payroll data shows. In the jobs most exposed to AI, employment for the youngest workers has dropped sharply since ChatGPT arrived, while everyone older in those same jobs kept right on growing. The second cut is harder to wave away. Take those same 22-to-25-year-olds and look at the jobs AI can barely touch, and their employment went up, not down. Each row is an age group. Within it, the dots show how employment has changed since November 2022, when ChatGPT launched: red for the most AI-exposed jobs, gray for the least-exposed. The wider the gap between the dots, the more a worker's exposure to AI seems to matter. For the youngest workers that gap is a chasm. For everyone else it nearly vanishes. A job's AI exposure is a research measure of how much of its everyday work an AI model can do or speed up. Software developers, customer service reps, and accountants are among the most exposed; hands-on roles like home health aides and the skilled trades are among the least exposed. The youngest workers are the only ones sliding In the most AI-exposed jobs, employment for 22-to-25-year-olds is down about 12% since ChatGPT launched, while every age group 31 and up grew. The next group, 26 to 30, is roughly flat. After that the line only points up: workers in their 30s, 40s, and 50s all gained ground in the exact same high-exposure jobs. The Stanford researchers who first flagged this have a clean explanation for why youth is the dividing line. Entry-level work leans on codified knowledge, the kind you can write down in rules and learn from a textbook, which is exactly what an AI model does well. Experienced workers lean on tacit knowledge, the judgment and context you only pick up on the job, which AI still fumbles. The junior tasks are the automatable ones, so the junior rungs are the ones being sawed off. It is the AI-exposed jobs, not a bad year to be young The same 22-to-25-year-olds grew about 7% in the least-exposed jobs, a swing of nearly 20 percentage points across the exposure scale. If this were just a rough market for young people in general, it would drag down the young everywhere. Instead it bites only where AI bites. It also does not look like mass firing. Both the Stanford team and the Dallas Fed, working from separate data, find the decline comes from a collapse in hiring rather than layoffs: young people are not being shown the door so much as never let in. Software development is the sharp end of it, where employment for developers aged 22 to 25 has fallen close to 20% from its late-2022 peak. The on-ramp is what broke, and it is the same on-ramp recent graduates have been struggling to find for a few years now. Experience is still a shield For workers 31 and older, AI exposure barely registers: the gap between the least- and most-exposed jobs is a few points at most. For 41-to-49-year-olds it actually tips the other way, with the AI-exposed jobs slightly ahead. That is the part that should reassure mid-career workers and worry new ones. The people who already cleared the entry-level gauntlet are insulated, because what they know is hard to hand to a model. The people trying to clear it now are competing with a tool that does the starter tasks for free. How much of this is actually AI? Here is where the honest version gets complicated, and the researchers are the first to say so. The same Stanford group later re-checked their own work and found the AI-exposed decline only turns statistically clean from 2024 onward; part of the earlier drop was probably something else, and they flatly state they do not think AI is the sole cause. Interest rates do not rescue the story either, since the exposed jobs are, if anything, less rate-sensitive than the rest. And not everyone is convinced there is a signal here at all. The Yale Budget Lab finds no clear link between AI exposure and employment or unemployment through August 2025, and a New York Fed study of job postings sees little sign of a distinct AI-driven drop in demand, noting the relative slide in AI-exposed roles started before ChatGPT shipped. The name researchers gave this pattern, canaries in the coal mine, is the right frame: a canary is an early warning, not a verdict. The youngest workers are where the effect would show up first if it is real, and right now they are the ones gasping. How this chart was made This chart was built by an AI agent and graded against the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: the Canaries dashboard from the Stanford Digital Economy Lab and ADP Research, built from anonymized ADP payroll records for about 25,000 firms. The Employment Index is set to 100 at November 2022 and broken out by age group and by occupational AI-exposure quintile; this latest release runs through April 2026. The cleaned dataset is available here. --- ## The Strait of Hormuz is Asia's oil lifeline, not the U.S.'s URL: https://www.randalolson.com/2026/06/21/strait-of-hormuz-oil-dependence/ Published: 2026-06-21 Categories: data visualization Tags: beautiful-charts-with-ai, Strait of Hormuz, oil, energy, geopolitics As the 2026 Strait of Hormuz crisis stoked U.S. oil fears, only about 7% of U.S. crude imports actually pass through it. Asia relies on it far more. Part of Teaching an AI Agent to Make Beautiful Charts When Iran threatened to close the Strait of Hormuz in June 2026, U.S. drivers braced for another gas-price shock and "Hormuz closed" raced up the trending charts. The United States, though, depends on the Strait less than almost any major economy. Roughly 20% of the world's oil passes through that single channel between Iran and Oman, and only a sliver of it is bound for the U.S. The world's most important oil chokepoint The Strait of Hormuz is the only sea route out of the Persian Gulf. About 20 million barrels of oil move through it every day, close to 20% of everything the world burns and more than 25% of all the oil that travels by ship, according to the U.S. Energy Information Administration. There is no easy detour: the pipelines built to skirt the Strait can carry only about 2.6 million barrels a day, a fraction of the flow. When this single waterway gets threatened, the entire oil market flinches. That is why a conflict thousands of miles from any U.S. coastline can still move the price at a gas station in Ohio. Why the U.S. barely feels it on supply The shale boom rewired where the U.S. gets its oil. The country now produces more crude oil than any nation in history, and in 2020 it became a net exporter of petroleum for the first time since at least 1949. The crude it does buy abroad comes overwhelmingly from next door: Canada sent the U.S. a record 4.1 million barrels a day in 2024, about 62% of the crude the country imports, almost all of it by pipeline, with Mexico and Latin America covering most of the rest. Only about 7% of U.S. crude imports come through the Strait of Hormuz, and Gulf oil now accounts for just 2% of all the petroleum the U.S. uses, the lowest level in nearly 40 years. The barrels that once tied the U.S. to the Persian Gulf have been replaced by domestic shale and a short pipeline ride from Alberta. Asia is the economy that actually lives on the Strait For Asia's big importers, Hormuz is not a headline risk, it is the supply line itself. Japan and South Korea have almost no oil of their own and buy nearly all of it abroad, most of it from the Gulf. In a normal year Japan draws close to 90% of its crude through the Strait and South Korea around 75%, while China and India each rely on it for roughly 45%. Japan leans on Hormuz about 12 times as heavily as the U.S. does. China is the Strait's single largest customer, taking about 37% of everything that crosses it, and together China, India, Japan, and South Korea account for 69% of all the crude that passes through. That dependence is exactly why those buyers scrambled the moment the Strait looked unsafe. As the 2026 conflict dragged on, Japan lined up crude that skips the Strait entirely, and India leaned harder on non-Hormuz suppliers. The chart shows the normal, peacetime picture. The war pushed Asia to do exactly what the U.S. never had to. The catch: prices are global even when the barrels aren't Oil is a global, fungible commodity, so a supply disruption anywhere lifts the price everywhere. Crude is the single biggest component of what U.S. drivers pay at the pump, so a Gulf scare still lands at the gas station even when the physical barrels never leave Asia. Before the worst of the crisis eased, the EIA expected Brent crude to peak near $115 a barrel and U.S. gasoline near $4.30 a gallon, and analysts warned a prolonged closure could push prices sharply higher. The U.S. is shielded from a supply cutoff, not from the price spike a Hormuz crisis sends around the world. Why the Strait has never actually closed For all the threats, the Strait of Hormuz has never been fully shut, not even during the 1980s "Tanker War," when Iran and Iraq launched more than 450 attacks on shipping and the oil kept moving. Part of the reason is that Iran needs the Strait too: nearly all of its own oil leaves through the same channel, and even during the 2026 conflict it kept exporting, sending about 11.7 million barrels to China through Hormuz. Closing Hormuz would choke Iran's own oil lifeline along with everyone else's, which is why the threat keeps coming back but the shutdown never quite arrives. The U.S. Navy's Fifth Fleet has guarded the waterway from its base in Bahrain for decades, and the world's spare bypass pipelines could move only a small fraction of what the Strait carries. So the cycle runs the same way each time: a threat, a price spike, a wave of headlines, and then the tankers keep sailing. The oil that crosses Hormuz was never really headed for the U.S. anyway. It is Asia that depends on this Strait, and Asia that has the most to lose on the rare day it ever truly closes. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It pulled the Strait of Hormuz flow data from the U.S. Energy Information Administration and the IEA, divided each economy's Hormuz crude by its total crude imports to get a dependence share, built the chart in Python, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: Strait of Hormuz crude flows from the U.S. Energy Information Administration (Today in Energy, drawing on Vortexa tanker tracking) and the IEA, with each economy's total crude oil imports from the JODI Oil World Database for 2024, and the European Union dependence figure from the Oxford Institute for Energy Studies. The dataset used for this chart is available here. --- ## The world's richest person has never been this far ahead of No. 2 URL: https://www.randalolson.com/2026/06/18/worlds-first-trillionaire-gap/ Published: 2026-06-18 Categories: data visualization Tags: beautiful-charts-with-ai, Elon Musk, wealth, billionaires, inequality Elon Musk became the world's first trillionaire on June 12, 2026. His fortune is now about 3.7x the second-richest person, the widest gap on record. Part of Teaching an AI Agent to Make Beautiful Charts On June 12, Elon Musk became the first person on Earth worth a trillion dollars. The number is hard enough to picture on its own, but the part that really stops you is the distance to everyone else. Musk is now worth more than the next 4 people on the global rich list combined, and for most of the past 25 years the person at No. 1 was barely ahead of No. 2 at all. Musk crossed a trillion dollars in a single afternoon The trigger was SpaceX going public. The company priced its shares at $135 and raised about $75 billion, the largest IPO in history, easily clearing the $25.6 billion Saudi Aramco raised in 2019. SpaceX began trading on June 12 at a valuation near $1.77 trillion, and Musk's slice of it was suddenly worth enough to push him past the line. That made him the world's first trillionaire, at roughly $1.1 trillion. The figure moves around by the day and by whose index you trust, but the milestone held: a single person's paper fortune is now bigger than the annual economic output of all but about 21 countries. The gap to No. 2 has never been this wide The chart tracks one simple thing: how many times richer the world's No. 1 is than the No. 2, on the Forbes list each year. For 2026 that ratio is about 3.7. The runner-up, Google co-founder Larry Page, is worth roughly $296 billion, which means the gap between first and second place is around $800 billion, larger than the entire GDP of Ireland. Put another way, Musk could lose hundreds of billions of dollars and still be the richest person alive. Nothing in the Forbes era looks remotely like it. The No. 1 spot has usually been a tight contest, not a blowout. For a generation, the race for No. 1 was a photo finish Look at the stretch from 2000 to 2024 and the dots barely lift off the floor. In a typical year the richest person was only about 1.15 times richer than the runner-up, and the lead almost never reached 1.5x. Whoever held the top spot, whether Gates, Bezos, or Arnault, had someone breathing down their neck. The closest call came in 2010, when Mexican telecom magnate Carlos Slim edged out Bill Gates for the top spot by just $500 million, $53.5 billion to $53.0 billion. That is 2 of the wealthiest people who have ever lived, separated by a rounding error. The old record was the dot-com bubble Before this year, the widest any No. 1 had ever pulled ahead came at the height of the late-1990s tech boom. On the 1999 Forbes list, Bill Gates was worth $90 billion to Warren Buffett's $36 billion, a ratio of exactly 2.5. Microsoft was near its peak and Gates briefly looked untouchable. Musk's 3.7x lead leaves the dot-com record far behind. Gates held that mark for more than 25 years. Everyone near the top got richer over those decades, but Musk got richer fast enough to pull away from the rest of the field. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It pulled the year-by-year No. 1 and No. 2 fortunes from the Forbes annual list, computed the ratio, built the chart in Python, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: net worth of the world's first- and second-richest people from the annual Forbes World's Billionaires list, with the live June 2026 figures from the Bloomberg Billionaires Index and cross-checked against reporting on the SpaceX IPO. The dataset used for this chart is available here. --- ## The 2026 NBA Finals drew the biggest TV audience since 1998 URL: https://www.randalolson.com/2026/06/17/nba-finals-viewership-2026/ Published: 2026-06-17 Categories: data visualization Tags: beautiful-charts-with-ai, NBA, basketball, sports, TV ratings The 2026 Knicks-Spurs Finals averaged 20.6 million viewers, double last year and the most since 1998, reversing a long slide in NBA TV ratings. Part of Teaching an AI Agent to Make Beautiful Charts For a couple of years now, the story about the NBA has been that nobody watches it anymore. Cord-cutting, load management, too many 3-pointers, a generation that would rather catch the highlights on a phone than sit through a broadcast. The ratings backed it up. Then the 2026 Finals doubled last year's audience and pulled the biggest crowd the league has had in front of a television since 1998. The decline was real The skeptics were not making it up. NBA Finals viewership peaked in 1998, when Michael Jordan won his sixth title with the Bulls and 29 million people a night tuned in. Game 6 of that series is still the most-watched NBA game ever, at 35.9 million. Nothing since has come close. From there it slid for 2 decades. By 2025, the Thunder-Pacers Finals drew 10.2 million viewers a night, roughly 1/3 of the 1998 audience. Some of that is the NBA's problem and some of it is everyone's: young viewers increasingly watch the league as short highlight clips rather than live broadcasts, and cable kept shedding the households that used to make a championship a default event. 2020 was the floor The low point came in the strangest season. The 2020 Finals drew just 7.5 million viewers a night, the smallest NBA Finals audience on record. The games were played in a sealed bubble in Orlando with no fans, pushed to October by the pandemic, and stacked against an election news cycle and Sunday night football. Almost everything that could pull eyeballs away was working at once. Then the Knicks doubled it The 2026 Finals reversed all of it. The New York Knicks won their first championship in 53 years, since 1973, and they did it in the country's largest media market against a San Antonio Spurs team built around Victor Wembanyama, the most hyped young player in the league. That is about as good as a TV matchup gets. The games delivered too. The Knicks erased a 29-point deficit in Game 4, the largest comeback in Finals history, and closed it out in Game 5 in front of 24.5 million viewers that peaked at 33 million at the buzzer. The series averaged 20.6 million viewers a night, double the 2025 number and the most since 1998. By the NBA's count, the series also racked up a record 15 billion video views on social media, nearly triple the prior year. The clips and the broadcast, it turns out, are not a zero-sum fight. Why 2026 still isn't 1998 Now, the honest asterisk. The viewership counts are not measured the same way across this whole chart. In late 2020 Nielsen started counting out-of-home viewing, the people watching in bars, gyms, and other households, and later expanded it to every market in the lower 48 states. Live sports gained the most from that change, so recent totals get a lift the 1998 figure never got. You can see it in the household ratings, which are measured more consistently. The 2026 Finals scored a 10.0 household rating against 1998's 18.7, a 47% drop, even though the raw viewer count was only 29% lower. So 2026 is not really back to Jordan-era reach. Strip out the measurement boost and it was probably still the best since 2017. Either way, a league that was supposedly dying just posted its biggest Finals in a generation, and a lot of people watched. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It compiled the year-by-year Nielsen viewership figures, built the chart in Python, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: average NBA Finals viewership per game, compiled from Nielsen figures published by Sports Media Watch and cross-checked against Sportico's reporting on the 2026 series and the 2020 to 2026 measurement change. The dataset used for this chart is available here. --- ## GLP-1 use peaks in middle age, then drops at 65 URL: https://www.randalolson.com/2026/06/15/glp1-use-peaks-middle-age/ Published: 2026-06-15 Categories: data visualization Tags: beautiful-charts-with-ai, GLP-1, Ozempic, public health, Medicare About 1 in 8 U.S. adults now take a GLP-1 like Ozempic, but use peaks in middle age: 22% at ages 50-64, falling to 9% at 65+ as Medicare stops covering it. Part of Teaching an AI Agent to Make Beautiful Charts The easy guess about GLP-1 drugs like Ozempic and Wegovy is that the people most likely to take them are the oldest and the heaviest: the group with the most diabetes and the most years of accumulated weight. The data says otherwise. Use of these drugs peaks in middle age, then falls off a cliff right at 65. About 1 in 8 U.S. adults now takes a GLP-1, but that average hides a sharp pattern by age. The peak is middle age More than 1 in 5 adults aged 50 to 64 (22%) currently take a GLP-1, against just 4% of adults under 30. That is more than 5 times the rate of the youngest adults, and nearly double the 12% average across all adults. Women are also more likely to be on them than men, 15% to 9%. Middle age is where the conditions these drugs treat, obesity and type 2 diabetes, become common, and where most people still have private insurance or the income to cover a drug that has listed at more than $1,000 a month. Use drops right at 65 The fall at 65 is not about health. Adults 65 and older have the highest diabetes rate of any age group, about 29%, yet their GLP-1 use drops to 9%, below the middle-aged group and below the national average. The reason is Medicare. A 2003 law bars Medicare's Part D drug benefit from covering any medication used for weight loss. Medicare will pay for a GLP-1 prescribed for diabetes or heart disease, the uses it treats as medical, but not for weight loss alone. So the weight-loss demand that drives the 50-64 peak largely disappears at 65, when private insurance gives way to Medicare. The number doubled in 18 months This is a recent surge. In May 2024, 6% of adults said they were currently taking a GLP-1. By late 2025 that had doubled to 12%. As many adults are on a GLP-1 today as had ever tried one a year and a half ago. Counting everyone who has used one at some point now puts ever-use at 18%. Price is still the wall Cost gates who gets them. Half of GLP-1 users told KFF the drugs are difficult to afford, and the brand-name versions have run past $1,000 a month at list price. Even at the middle-age peak, price holds use well below the number of people who would medically qualify. The cliff at 65 is about to move That cliff is about to get smaller. After scrapping a Biden-era plan to cover obesity drugs in early 2025, the Trump administration reversed course: it struck deals with Novo Nordisk and Eli Lilly in November 2025 to cut GLP-1 prices, and a voluntary Medicare demonstration begins July 1, 2026 that lets participating drug plans cover GLP-1s for weight loss with a $50 monthly copay. If it holds, the right end of this chart could look very different a year from now. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It pulled the age breakdown from the KFF Health Tracking Poll, built the chart in Python, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: the KFF Health Tracking Poll, fielded October 27 to November 2, 2025 among 1,350 U.S. adults (margin of error +/- 3 percentage points). Current GLP-1 use by age is reported in KFF's writeup of the poll. The dataset used for this chart is available here. --- ## ICE detention hit a record, and most detainees have no conviction URL: https://www.randalolson.com/2026/06/12/ice-detention-no-criminal-conviction/ Published: 2026-06-12 Updated: 2026-06-20 Categories: data visualization Tags: beautiful-charts-with-ai, immigration, ICE, immigration enforcement, criminal justice U.S. immigration detention hit a record 70,000+ people in early 2026, and TRAC data shows about 7 in 10 detainees have no criminal conviction. Part of Teaching an AI Agent to Make Beautiful Charts In June 2026, Congress handed U.S. Immigration and Customs Enforcement and the Border Patrol roughly $70 billion in new money, the second multibillion-dollar infusion in a year. The case for it has been steady: the government says it is going after "the worst of the worst," the criminal aliens it describes in a near-daily stream of press releases. The people actually sitting in ICE detention look little like that description. On the most recent count, 42,722 of the 60,311 people in custody, about 7 in 10, had no criminal conviction at all. And this is not something the crackdown invented. It is what ICE detention has looked like for years. The chart splits everyone held in ICE custody into two groups: people convicted of a crime, and people with no conviction. The total rises and falls with how many ICE is holding at any given time. A record number of people are in ICE detention In late January 2026, ICE held 70,766 people, the most it had ever held at once and the first time the agency topped 70,000. The detained population had roughly doubled off its 2024 lows before easing back to about 60,000 by April. That growth was bankrolled. A $75 billion windfall in the summer of 2025 made ICE the highest-funded law enforcement agency in the federal government and let it double its ranks in a matter of months. The June 2026 package, about $38 billion of it earmarked for ICE, locked in that footing through the rest of the president's term. Almost all of the growth has been people with no conviction About 3 of every 4 people added to detention during the 2025 buildup had no criminal conviction. The number of convicted people in custody did rise too, but far more slowly, and they have stayed a clear minority of everyone held. TRAC's own accounting is starker over the sharpest stretch. Between September and November 2025, detention grew by 5,373 people, and 5,209 of them, about 97%, had no criminal conviction. The same shift shows up at the arrest stage: the share of people ICE arrested who had no criminal record at all climbed from about 22% in early 2025 to roughly 40% by that fall. It was never mostly criminals, even before the crackdown Even in early 2024, under the prior administration, about 3 in 4 ICE detainees had no criminal conviction. The share has sat between about 69% and 76% across the entire span the chart covers. That is the part the "worst of the worst" framing skips. The no-conviction share is not a number the 2025 surge drove up. It is a baseline that held steady while the system around it more than doubled in size. What "no criminal conviction" does and does not mean "No criminal conviction" is not the same as "no record." It includes people with pending charges who have not been convicted of anything. Strip those out and the picture is still striking: as of early January 2026, about 48% of detainees had neither a conviction nor a pending charge, and by April that share was about 40%. The convictions that do exist are not all serious, either. TRAC notes that many of those counted as convicted committed only minor offenses, including traffic violations. When PolitiFact checked the claim that roughly 70% of detainees had been convicted of or charged with a crime, it did not hold up: even counting pending charges, closer to half had neither. None of this says no one in ICE custody is dangerous. It says the detained population is mostly people the system has not convicted of anything, and that this was true before the latest expansion and stayed true through it. How this chart was made This chart was built by an AI agent and graded against the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: Immigration Detention Quick Facts and the biweekly detention series from TRAC Immigration at Syracuse University, based on ICE detention data. "No criminal conviction" combines TRAC's pending-charges and other-immigration-violator categories, so it includes people with charges that have not led to a conviction. The two-week snapshots run from January 2024 through April 4, 2026, with a Sep-Nov 2025 gap that reflects data ICE withheld during the federal shutdown. The cleaned dataset is available here. --- ## U.S. abortions rose after Dobbs, driven by mailed pills URL: https://www.randalolson.com/2026/06/12/us-abortions-rose-after-dobbs/ Published: 2026-06-12 Updated: 2026-06-20 Categories: data visualization Tags: beautiful-charts-with-ai, abortion, telehealth, public health, wecount Dobbs was supposed to cut U.S. abortions. It did not. The national total rose after 2022, and the whole increase is telehealth: abortion pills by mail. Part of Teaching an AI Agent to Make Beautiful Charts The widely shared story about the Dobbs decision is simple: the Supreme Court let states ban abortion, bans spread across the South and Midwest, so abortions fell. The first half is right. The second half is not. Since the Court overturned Roe in June 2022, the number of abortions in the U.S. has gone up, not down. If the bans had cut the number of abortions, the line would fall. It doesn't. The national count rose roughly 18% to its early-2025 peak, and one thing accounts for the entire increase: abortion pills in the mail. The national total went up, not down Spring 2022, the last full quarter mostly before the ruling, saw about 253,750 abortions in the U.S. health care system. By early 2025 the quarterly figure peaked at 300,250. The monthly average climbed every year since: 79,600 in 2022, 88,200 in 2023, 95,300 in 2024, and 98,800 in the first half of 2025. January 2025 was the single biggest month on record for this count, with 107,740 abortions. This is not a single tracker's quirk. Guttmacher, counting on a different method, found more than 1 million abortions in the formal health care system in 2023, the most since 2012. Two independent efforts that disagree on the exact numbers agree on the direction: up. The whole increase is pills in the mail Split the total into in-person care and telehealth and the mechanism is obvious. In-person care, the lower band, actually fell about 12%, from 242,110 abortions a quarter to 211,930. Telehealth, the band stacked on top, went from 5% of all abortions to 27%, growing nearly 7 times over in absolute terms. That single band added about 68,000 abortions a quarter, more than enough to cover the in-person decline and push the total higher. Telehealth here means a clinician consults with a patient remotely and mails mifepristone and misoprostol, the drugs used for a medication abortion. Medication abortion is now the majority of U.S. abortions, about 63% in 2023. The rise after Dobbs did not come from more clinics. It came from the mailbox. Shield laws carried the pills into ban states The pills increasingly go to the places that banned the procedure. 13 states enforce total bans, and in those states nearly all the abortions still happening are telehealth pills mailed in from providers in other states. Those providers work under shield laws, state protections that a handful of states first enacted in 2023. By June 2025 shield-law providers were mailing roughly 14,770 abortions a month. One honest caveat lives at the mid-2023 marker on the chart: that is when shield-law provision began and the #WeCount project broadened its data collection, so part of the 2023 jump is better measurement rather than pure growth. The trend holds without it, since telehealth kept climbing for 2 more years. Travel did the rest. More than 169,000 people, about 1 in 6, crossed state lines for an abortion in 2023, double the share in 2020. Bans changed how and where people got abortions, not whether. What the count leaves out This chart only counts abortions provided by licensed clinicians. It leaves out self-managed abortions, the pills people get outside the formal system through networks like Aid Access or overseas pharmacies, which jumped by tens of thousands of requests in the months after Dobbs. The 2022 and 2025 windows are also partial years, and #WeCount imputes 19% to 28% of its counts. Every one of those gaps points the same way: the true national total is higher than the chart shows, not lower. The count is a floor. The fight moved to the pills Because the pills are now the story, the legal fight is too. In May 2026 the Supreme Court left mifepristone available by mail and telehealth while Louisiana v. FDA moves through the courts, and the FDA moved ahead with a new safety review of a drug it approved 25 years ago. States are fighting over the shield laws directly: Texas fined a New York doctor more than $100,000 for mailing pills into the state, Louisiana indicted her, and New York refused to extradite her. Abortion is back on the 2026 ballot in Missouri, Nevada, and Virginia. The data does not tell anyone which side is right, and it is not meant to. It says the simplest version of the post-Dobbs story, the one where bans cut the number of abortions, is not what the numbers show. Whatever the bans did, they did not lower the count. They moved it into the mail. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It pulled the quarterly figures from the #WeCount report, built the chart in Python, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: the Society of Family Planning's #WeCount Report, April 2022 to June 2025 (released December 9, 2025), which counts clinician-provided abortions including telehealth and shield-law provision and excludes self-managed abortions. The quarterly dataset used for this chart is available here. --- ## Curaçao is the smallest nation ever at a World Cup URL: https://www.randalolson.com/2026/06/10/world-cup-2026-population-gap/ Published: 2026-06-10 Categories: data visualization Tags: beautiful-charts-with-ai, world cup, fifa, soccer, curacao The 48 nations at the 2026 World Cup span a 2,181x gap in population, from the United States (340M) to Curaçao (156,000), the smallest ever to qualify. Part of Teaching an AI Agent to Make Beautiful Charts The 2026 World Cup kicks off on June 11 with the most lopsided field the tournament has ever assembled. The United States, the biggest of the 3 hosts, has about 340 million people. Curaçao, a Caribbean island in the same 48-team draw, has roughly 156,000. Curaçao's entire population would be a rounding error in the U.S. census. Sort all 48 nations by population and the spread runs 2,181 to 1, top to bottom. That is hard to picture, so I plotted every nation in the field on a single scale. Curaçao is the smallest nation ever to reach a World Cup At about 156,000 people, Curaçao is the smallest country, by population and by land area, ever to make a men's World Cup. It takes the record from Iceland, which had around 350,000 when it reached the 2018 finals. The old record holder was more than twice Curaçao's size. How an island of 156,000 pulls this off is its own story. Curaçao is a constituent country of the Kingdom of the Netherlands, and it leans hard on that link: all but one of the 26-man squad was born in the Netherlands, and 16 of them played for Dutch youth teams. Coached by veteran Dutchman Dick Advocaat, who will become the oldest manager in World Cup history, they qualified unbeaten and clinched it with a nervy 0-0 draw away at Jamaica in November. A 48-team field changed who gets in None of this happens in the old format. 2026 is the first World Cup with 48 teams instead of 32, the biggest expansion in the tournament's history. FIFA approved it in 2017 and pitched it as a way to make the game "truly global" and give more of its 211 member associations a realistic shot. The new places did not fall evenly. Africa sent 5 nations to the 2022 World Cup and sends 10 this time; Asia rose from 6 to 9; Europe, already the biggest bloc, went from 13 to 16. The expansion was, in practice, an opening for smaller and less established football nations, which is how a single field can hold both the United States and Curaçao. Cape Verde is the only other nation under a million Curaçao is not alone at the bottom. Cape Verde, an Atlantic archipelago off West Africa with about 525,000 people, is the only other nation in the field under a million, and like Curaçao it is making its World Cup debut. It is the third-smallest country ever to qualify, behind Curaçao and Iceland. The "Blue Sharks" earned it on the field, beating Eswatini 3-0 in October to top their African group 4 points clear of Cameroon. 2 of the 4 debutants this year, Curaçao and Cape Verde, rank among the 3 smallest nations ever to reach the tournament. The other 2 newcomers, Jordan and Uzbekistan, are far bigger. The giants at the top, and the limits of being big At the other end sit the heavyweights. 6 nations in the field have more than 100 million people: the United States, Brazil, Mexico, Japan, Egypt, and DR Congo. The median nation has about 30 million, and the 48 countries hold roughly 2.2 billion people between them, about 27% of everyone alive. Population, though, does not buy a spot. The 2 most populous countries on Earth, India and China, with about 1.4 billion people each, did not qualify. A Caribbean island of 156,000 did. A bigger talent pool helps, but it has never been the thing that gets a team to the World Cup. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It researched the data, built the chart in Python, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: Population figures are from the World Bank (2024), except England and Scotland, which the World Bank counts only inside the United Kingdom; those come from the Office for National Statistics and National Records of Scotland (mid-2023). The 48-team field is per the 2026 FIFA World Cup record. The dataset used for this chart is available here. --- ## The rise in LGBTQ+ identification is mostly young women URL: https://www.randalolson.com/2026/06/09/lgbtq-identification-by-generation/ Published: 2026-06-09 Updated: 2026-06-20 Categories: data visualization Tags: beautiful-charts-with-ai, lgbtq, demographics, gallup, gen z Nearly 1 in 3 Gen Z women in the U.S. identifies as LGBTQ+, versus 12% of Gen Z men. Gallup data shows the generational surge is mostly about women. Part of Teaching an AI Agent to Make Beautiful Charts You have probably seen the headline: Gen Z is far more likely to identify as LGBTQ+ than any generation before it. That much is true. But it flattens a more specific story. The rise is not spread evenly across young people. It is concentrated among young women, and most of it runs through bisexuality, often quietly enough that the people around them might not know. Gallup has tracked this in the U.S. since 2012. Its 2024 survey of more than 14,000 adults is the most recent one that splits identification by both generation and gender, and the split is where the real story lives. The chart is really a timeline of acceptance Read it from the bottom up and you are watching a country change its mind. The Silent Generation came of age when homosexuality was officially classified as a mental illness, a label that stood until 1973. Gen Z came of age after the Supreme Court made same-sex marriage a settled right. Almost nobody in the oldest group says they are LGBTQ+. Nearly 1 in 4 of the youngest does. That climb tracks the country's own turn. Support for same-sex marriage went from 27% in 1996 to 70% by 2021. As acceptance rose, the cost of saying it out loud dropped, and more people said it. The line is steep because the change was fast, and each generation grew up in a more accepting version of the same country. It's young women, not just young people Split the youngest generations by gender and the tidy "kids these days" framing falls apart. Among Gen Z, 31% of women identify as LGBTQ+, against 12% of men. The same gap shows up for millennials and fades with age. What everyone files under youth is, to a large degree, a shift among young women. The honest answer to why is that nobody is sure yet. Psychologist Lisa Diamond's research argues that women's sexuality tends to be more fluid over a lifetime, and Pew finds 16% of women under 30 call themselves bisexual, against 5% of men. I went in hoping for a clean explanation and came out without one, which is its own kind of finding. Most of it is bisexuality, and most of it is quiet There is another layer under the gender gap. The label doing most of the work is bisexual. Among LGBTQ+ members of Gen Z, most are bisexual; among the oldest LGBTQ+ adults, almost none are. Older generations are mostly gay and lesbian. Younger ones are mostly bi. Bisexuality is also the least visible part of the spectrum. Most bisexual adults with a partner are with someone of a different sex, and most are not out to the people who matter to them. So a lot of this rise is not a parade. It is often a woman in a relationship with a man, answering a survey question differently than she would have a decade ago. The gap is a forecast The overall number recently stopped rising. It has held near 9% for 2 years, which Gallup calls essentially unchanged. It is tempting to read that as the trend topping out. But a generation gap this wide is less a snapshot than a forecast. As the Silent Generation gives way to Gen Z, the national share keeps climbing on its own, with nobody changing their mind. The country's LGBTQ+ population is growing less because adults are rethinking who they are and more because of who is aging out and who is coming up. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It researched the data, built the chart in Python, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: Gallup's annual LGBTQ+ identification survey, specifically the 2024 release (the most recent year Gallup breaks identification out by both generation and gender) and the 2025 update. The dataset used for this chart is available here. --- ## Honey bees are dying at record rates, but the colony count holds steady URL: https://www.randalolson.com/2026/06/07/us-honey-bee-colony-paradox/ Published: 2026-06-07 Categories: data visualization Tags: beautiful-charts-with-ai, honey bees, agriculture, bees, colony collapse U.S. beekeepers lost a record 55.6% of their colonies in 2024-25, yet the total number of managed honey bee colonies has barely moved in 15 years. Part of Teaching an AI Agent to Make Beautiful Charts U.S. beekeepers just had the worst year of colony losses ever recorded. From April 2024 to April 2025 they lost an estimated 55.6% of their managed honey bee colonies, more than in any year since the survey started tracking annual losses. And yet the total number of honey bee colonies in the country is right about where it has been for the past 15 years. Those two facts seem to contradict each other. Here is how they fit together. Beekeepers just had their worst year on record An estimated 55.6% of managed U.S. colonies died between April 2024 and April 2025, the highest annual loss since the national survey began in 2010-11. The winter portion alone was 40.2%, beating the previous winter record of 37.7% from 2018-19, according to the Auburn University and Apiary Inspectors of America survey. The leading suspect is not a mystery. USDA scientists traced the early-2025 collapses to viruses spread by the Varroa destructor mite, and found mites resistant to amitraz, the main miticide, in nearly every sample. Project Apis m. put the economic hit above $600 million. This was a brutal year by any measure. Yet the number of colonies barely moved Here is the part that surprised me. A 55.6% loss sounds like a population in freefall, but the headline colony count did not fall off a cliff. USDA pegs the number of managed honey-producing colonies at 2.41 million in 2025, down about 10% from 2010 and still sitting inside the same 2.4 to 2.8 million band it has held since the late 2000s. The record death rate was not an anomaly either: since 2011, beekeepers have lost an average of 42% of their colonies every single year. Even in their best recent year, they still lost nearly 30%. Losses at this scale have been the normal backdrop for over a decade, and the count held steady through all of it. Honey bees are livestock, and beekeepers restock the herd The trick is that "loss rate" measures how many colonies die, not how many exist at year's end. Honey bees are managed livestock, not wildlife. When colonies die, beekeepers rebuild: they split a strong surviving hive into 2 or 3, buy new queens, and order packaged bees, and a split can be back to working strength in about 6 weeks. The USDA's Economic Research Service has made this point for years, in a report bluntly titled "Despite Elevated Loss Rate Since 2006, U.S. Honey Bee Colony Numbers Are Stable." That rebuilding is doing enormous, invisible work. If beekeepers replaced none of those dead colonies, compounding the measured loss rates would have left under 1,000 colonies by 2025 instead of 2.4 million. The flat line in the bottom chart is not the absence of death. It is beekeepers running up the down escalator, fast enough to stay in place. The real collapse already happened, decades ago Step back far enough and the bee crash is a 20th-century story, not a current one. The U.S. had about 5.9 million honey-producing colonies in 1947 and roughly 2.4 million by 2008. That long slide was driven by cheap imported honey undercutting the economics of beekeeping, lost forage, fewer beekeepers, and the arrival of tracheal mites in the 1980s and Varroa in 1987. Colony Collapse Disorder, the scary name that launched a thousand headlines, was coined in 2007 and largely faded as the dominant cause after about 2010. The bees stopped vanishing in that specific, mysterious way. They kept dying at high rates from causes we mostly understand now: mites, the viruses they carry, pesticides, and poor nutrition. The number of colonies, meanwhile, has been flat to rising the entire time. The bees in real trouble are the wild ones This is where the popular "save the bees" story goes sideways. The honey bee is a managed, non-native agricultural animal, closer to a chicken than to an endangered species, and its numbers are fine. The bees that are actually declining are wild and native ones. The rusty patched bumble bee became the first bumble bee in the U.S. listed as endangered, back in 2017, after disappearing from most of its range, and many other native bees are losing ground to habitat loss and pesticides. There is even growing evidence that packing huge numbers of managed honey bees into one place can crowd out wild bees competing for the same flowers, though the size of that effect is still debated. Honey bees are not going extinct. We keep restocking them by the millions, in part because the vast majority of the country's colonies get trucked to California every February to pollinate a single crop: almonds. The harder problem, and the quieter one, is the bees nobody is keeping. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It researched the data, built the chart in Python, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: colony counts are from USDA's National Agricultural Statistics Service annual Honey report (honey-producing colonies). Annual loss rates are from the Apiary Inspectors of America and Project Apis m. survey for 2023-24 and 2024-25, and from the Bee Informed Partnership national survey for earlier years. The dataset used for this chart is available here. --- ## CO2 is piling up faster every decade, and 2026 just set a record URL: https://www.randalolson.com/2026/06/05/co2-rising-faster-every-decade/ Published: 2026-06-05 Categories: data visualization Tags: beautiful-charts-with-ai, climate change, carbon dioxide, co2, environment Atmospheric CO2 hit a record 432 ppm in May 2026, but the deeper story is the pace: CO2 is now piling up about 3 times faster than it did in the 1960s. Part of Teaching an AI Agent to Make Beautiful Charts Carbon dioxide is the main driver of climate change. It traps heat, so the more of it in the air, the warmer the planet runs, which is why its concentration is the most closely watched number in climate science. Almost everyone has seen the chart of that number climbing since the 1950s. But the line is not just going up, it is getting steeper. The air took on more CO2 in the last decade than in any 10-year stretch since record-keeping began in 1959, and this past May the monthly average at Mauna Loa hit 432 parts per million, the highest ever measured. This is the famous Keeling Curve, named for Charles David Keeling, who started the measurements on a Hawaiian volcano in 1958. Each point is the average concentration of CO2 in the air that year. Parts per million is literal: at 432 ppm, 432 of every million molecules of air are CO2. That share sounds trivial, but CO2 traps heat so well that this trace amount helps set the planet's temperature, and pushing it higher pushes the planet warmer. More CO2 every year than the year before Every single year since 1959, there has been more CO2 in the air than the year before. The annual average has climbed from 316 ppm in 1959 to 427 ppm in 2025, and the seasonal peak this past May reached 432 ppm. Not one year reversed the trend. The dashed line marks the pre-industrial level of about 280 ppm, where CO2 sat for thousands of years before we started burning coal and oil. We are now more than 50% above that. Across the 800,000-year ice-core record, CO2 never once climbed past about 300 ppm, according to NOAA. The last time the air held this much carbon dioxide was roughly 3 million years ago, in the mid-Pliocene, when the planet was 2.5 to 4 degrees Celsius warmer and seas were far higher. The level is only half the story The pace is the other half, and it hides inside that smooth line, because "how fast CO2 is rising" is a step removed from "how much is up there." Here is the same Mauna Loa data, cut a different way. Each line takes a decade, resets it to zero in its first year, and stacks them on top of each other. A steeper line means CO2 piled on faster that decade. The decades fan out in almost perfect order, slowest at the bottom and steepest at the top. That spread is the acceleration. Each decade outpaces the last CO2 now builds up about 3 times faster than it did in the 1960s. Back then the air gained roughly 0.88 ppm a year. The 2010s ran at 2.41 ppm a year, and the 2020s are running hotter still at 2.63. In total, the 2010s added about 24 ppm, more than any other decade on record. The reason is blunt: we keep emitting more. Global fossil-fuel CO2 emissions hit a record 37.4 billion tonnes in 2024 and rose again, with the Global Carbon Project reporting no sign of a peak. As long as yearly emissions keep setting records, the amount in the air climbs faster and faster. 2024 was the biggest single-year jump on record From 2023 to 2024, CO2 at Mauna Loa jumped 3.33 ppm, the largest single-year increase in the record, and NOAA's global figure of 3.75 ppm set a record too. Record emissions set the baseline, and the 2023 to 2024 El Nino pushed it higher. Hot, dry conditions across the tropics gutted the land carbon sink, the forests and soils that normally soak up part of what we emit, so more of it stayed in the air. As that El Nino faded, the 2025 increase eased back to 2.23 ppm, still high by any historical standard. The decade that broke the pattern The fan has a single hiccup: the 1990s came in a touch slower than the 1980s. That dip traces almost entirely to a single volcano. When Mount Pinatubo erupted in June 1991, it threw enough sulfur into the stratosphere to cool the planet by about 0.5 degrees Celsius, which let forests and oceans pull down extra carbon for a couple of years; the global growth rate briefly fell to around 0.6 ppm in 1992, per research in Geophysical Research Letters. The collapse of the Soviet Union around the same time cut its emissions sharply and helped hold the decade down. Then the climb resumed, faster than before. The pace is the point It is tempting to fixate on each new record ppm number, and a fresh one lands almost every May. But the records are a symptom. What matters is that the line keeps bending upward, decade after decade, because we keep adding more carbon than the year before. At the current pace, the Global Carbon Project puts the world on track to blow past the 1.5 degree Celsius warming limit within a few years. The good news, if there is any, is that this is the most direct dial we have. The slope of that fan is set by how much we emit. Bend emissions down and the lines flatten. Nothing in 67 years of data says that is impossible. It just has not happened yet. How this chart was made These charts were built by an AI agent and graded against the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: Annual mean and monthly atmospheric CO2 from the NOAA Global Monitoring Laboratory, measured at the Mauna Loa Observatory in Hawaii (record begins 1959). Because of a 2022 eruption at Mauna Loa, NOAA's measurements briefly came from nearby Maunakea between late 2022 and mid-2023. The cleaned annual-average dataset is available here, and the per-decade data behind the second chart is here. --- ## New U.S. college grads now have higher unemployment than the average worker URL: https://www.randalolson.com/2026/06/04/recent-grad-unemployment-flip/ Published: 2026-06-04 Updated: 2026-06-20 Categories: data visualization Tags: beautiful-charts-with-ai, college graduates, unemployment, labor market, higher education A recent U.S. college degree long meant lower unemployment than the average worker. That edge flipped in 2019, and the gap is now the widest on record. Part of Teaching an AI Agent to Make Beautiful Charts A fresh college degree used to come with a quiet edge in the job market. New grads had better odds of landing work than the average worker, and that edge held for as long as anyone tracked it. Not anymore. They now face higher unemployment than the workforce as a whole, and the gap is the widest on record. What makes this strange is the timing. The reversal did not start with ChatGPT, and it did not start with the pandemic. It started in early 2019, before either one was on the radar. The chart tracks a single number, a recent grad's unemployment rate minus the rate for all workers. Below the zero line grads come out ahead of the typical worker, and above it they fall behind. The comparison is worth pinning down. "All workers" is the whole U.S. labor force, and most of them are older and more experienced than a new graduate, so a fresh grad starts at a natural disadvantage. For decades the degree more than canceled that disadvantage out. Now it does not. For decades, the degree was a buffer A recent grad almost always had a better shot at being employed than the average worker. The cushion was real, and it was biggest exactly when the economy was worst. The edge peaked in the depths of the Great Recession. In mid-2010, grads were around 7% unemployment while the workforce overall ran close to 10%, the widest that advantage ever got. Recessions gut construction and manufacturing first, sectors that lean heavily on workers without degrees, so a diploma was worth the most precisely when jobs were vanishing. The edge vanished in 2019, before AI and before COVID In February 2019 the gap crossed zero, and the 12-month average has stayed positive every month since. That timing rules out both of the easy explanations. The flip predates the generative-AI boom by years, and COVID by a year. This was a slow structural drift, not a sudden shock. The Cleveland Fed traces the erosion back further still. The job-finding advantage of young grads has been fading since around 2000, and their edge over high-school workers closed around 2019. The pandemic did not cause it either, and the 2020 spike is the clearest evidence. When unemployment exploded that year, both lines shot up together, so the gap held roughly steady through 2020 and 2021. The penalty was already there. The lockdowns just buried it under a bigger number. Now it is the widest gap on record By early 2026 recent grads sat at 5.6% unemployment against 4.2% for all workers, the widest gap on record. The spread has grown in almost every year since the 2019 flip. What makes the record stranger is the backdrop. This is not a recession story. Overall unemployment sits at a healthy 4.2%, yet new grads are the ones struggling. Every earlier spike in their unemployment arrived with a broad downturn. This one is theirs alone. Unemployment is only half the picture. Of the new grads who do have jobs, about 41% are underemployed, working roles that never required a degree in the first place. Remote work, or AI? So what broke? The honest answer is that economists are still arguing about it. In June 2026 the New York Fed made the case that remote work, not AI, is the main culprit, pinning about 64% of the rise in young-grad unemployment on it. Employers, the Fed argues, are wary of hiring inexperienced people into remote roles, where the on-the-job mentorship that turns a new grad into a productive worker is hard to deliver. The timing fits, since the climb started well before AI took hold. Stanford researchers see AI's fingerprints anyway. Their study found that early-career workers ages 22 to 25 in the most AI-exposed jobs saw employment fall about 16% since late 2022, a drop that survived even after they stripped out remote-friendly roles. Both can be true. Either way, the entry-level rungs are the ones being pulled out, and tech is the sharpest edge. Recent computer science graduates now post some of the highest unemployment rates of any major, after the number of CS degrees more than doubled into a shrinking pile of openings. The on-ramp broke, not the degree This is an entry-level problem, not proof that a degree stopped paying off. Older degree-holders are doing fine. U.S. workers 25 and up with a bachelor's or more had just 2.8% unemployment in April 2026, per the Bureau of Labor Statistics, comfortably below the rate for high-school grads. The damage is concentrated almost entirely in the young. Since 2019, recent grads have taken the brunt of the rise while unemployment for older degree-holders has barely moved, per the St. Louis Fed, and the New York Fed still pegs the lifetime return on a degree near 12.5%. New grads have not fallen behind their peers who skipped college, either. Young workers without a degree sit at 7.2% unemployment, well above the grads' 5.6%. A degree still beats no degree. What it no longer does is beat the average. None of this is settled. The Economic Policy Institute argues the picture is more mixed, with the college wage premium flat for years and new grads still doing no worse than young workers without degrees. The degree still opens the door. It just no longer gets you through it faster than everyone else. How this chart was made This chart was built by an AI agent and graded against the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: The Labor Market for Recent College Graduates from the Federal Reserve Bank of New York, built from the U.S. Census Bureau and Bureau of Labor Statistics Current Population Survey. "Recent graduates" are nonstudents ages 22 to 27 with at least a bachelor's degree, "young workers" are ages 22 to 27 without one, and "all workers" are ages 16 to 65. The cleaned dataset is available here. --- ## Gas prices feel like a record in 2026. Adjusted for inflation, they are not URL: https://www.randalolson.com/2026/06/03/us-gas-prices-inflation-adjusted-1973-2026/ Published: 2026-06-03 Updated: 2026-06-20 Categories: data visualization Tags: beautiful-charts-with-ai, gas prices, inflation, energy, oil prices U.S. gas prices jumped nearly 60% since January in the 2026 oil crisis, but adjusted for inflation they sit below the 1981, 2008, and 2022 peaks. Part of Teaching an AI Agent to Make Beautiful Charts Gas is the most expensive it has been in years, the war headlines are everywhere, and "record gas prices" is back in everyone's feed. The pump number stings. But adjust for inflation and the "record" evaporates: U.S. gas cost more in 1981, more in 2008, and more in 2022 than it does today. The number on the pump is near an all-time high. What a gallon actually costs you, in real money, is nowhere near it. The red line is what a gallon actually costs once every year is put in today's dollars. The gray line is the price you literally hand over at the pump, near its all-time high. The space between them is inflation, and it is the whole story. The fastest climb since 2022 The run-up is real. U.S. regular gas went from $2.81 in January 2026 to $4.48 in May, a 59% jump in 4 months. The trigger was the war between Iran and a U.S.-Israel coalition, and Iran's closure of the Strait of Hormuz, the chokepoint that carries about 20% of the world's oil. Prices have since eased to around $4.30 as a fragile ceasefire holds. A jump that fast in a single spring is the kind of move that makes people reach for the word "record." It just is not one. 2008 is still the one to beat The real all-time high was June 2008, at $6.20 a gallon in today's money. Crude oil hit $147.27 a barrel on July 11, 2008, pushed there by surging Chinese demand, a weak dollar, and heavy financial speculation. The pump price peaked just over $4 a gallon, which sounds close to today until you account for 2 decades of inflation. That is $1.72 a gallon more than spring 2026. 1981 was worse too 1981 held the record before 2008 took it. In March 1981, regular gas hit $5.32 a gallon in today's dollars. The Iranian Revolution had cut Iran's output from 6 million barrels a day to about 1.5 million, the squeeze behind the 1979 gas lines, and the Iran-Iraq war that followed kept prices climbing. For 27 years, no oil shock topped it. Even 2022 cost more The most recent peak is the one people remember best. When Russia invaded Ukraine, the national average set an all-time record just above $5.00 a gallon in June 2022, per AAA. That was the highest price ever at the pump in plain dollars. Adjust it for inflation and it works out to about $5.55, still above spring 2026 and still short of 2008. Expensive, just not historic None of this means gas is cheap. Today's price is higher, in real terms, than 83% of the months since 1973. Only 14 of the last 54 years had even a single month that cost more. The early 2010s are the clearest case: gas held above where it sits today for nearly 4 years straight, from 2011 into 2014, a slow grind rather than a spike. So both things are true at once: gas really is expensive right now, and it is nowhere near a record. Even if the pump reaches $5 this summer, as some forecasters warn, that would land around the 1981 line, not the 2008 one. How this chart was made This chart was built by an AI agent and graded against the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: Retail gasoline prices from the U.S. Energy Information Administration (national regular gasoline, with the pre-1990 tail from the EIA Monthly Energy Review, Table 9.4), adjusted for inflation with the Bureau of Labor Statistics CPI-U. The cleaned dataset is available here. --- ## The countries that work the most hours produce the least per hour URL: https://www.randalolson.com/2026/06/02/oecd-hours-worked-vs-productivity/ Published: 2026-06-02 Categories: data visualization Tags: beautiful-charts-with-ai, labor productivity, working hours, oecd, economics Across 40 OECD economies, the countries that log the most annual hours have the lowest output per hour. Why working longer rarely means producing more. Part of Teaching an AI Agent to Make Beautiful Charts The average worker in Germany puts in about 1,330 hours a year, fewer than in any country the OECD tracks. The average worker in Mexico puts in about 2,200, near the top of the list and roughly 2/3 more time on the job. Common sense says the country logging all those extra hours should be out front. It is not even close, and it runs the other way: a German hour of work generates about $98 of output, a Mexican hour about $25. About 1/4 as much, in exchange for all that extra time. Germany and Mexico are not a fluke. Put all 40 OECD economies on the same axes, hours worked against output per hour, and the whole field leans the same way. The longer a country works, the less it tends to get out of each hour. More hours, less output per hour The slope is steep and the fit is tight, a correlation of -0.74. Every extra 100 hours a country works in a year lines up with about $10 less output per hour. The cleanest version of the story lives at the extremes: Germany, Denmark, Norway, and the Netherlands sit in the top-left with short hours and high output, while Mexico, Colombia, and Costa Rica anchor the bottom-right with the longest hours and the lowest. One reason hours fail to buy output is that hours are not a clean input. They hit diminishing returns, and the dropoff is sharper than most people assume. Stanford economist John Pencavel went through the output records of First World War munitions workers and found that productivity per hour holds steady up to about 48 hours a week, then falls off a cliff. Past 56 hours, extra time bought almost nothing: a 70-hour week produced about as much as a 56-hour one. The back half of a long week is already low-yield before capital or technology enters the picture. Hours run into diminishing returns, and economies hit that wall long before they stop piling on hours. Where people work least and produce most The short-hours cluster did not get there by slacking off. The Netherlands has the highest rate of part-time work in the OECD, with close to 1/2 of working Dutch women in part-time roles, which pulls the average hours per worker down while output per hour stays near $100. Germany works the fewest hours of anyone, and that is negotiated, not accidental. Sectoral union contracts set the workweek below the legal ceiling, and in 2018 IG Metall won workers the right to drop to a 28-hour week for up to 2 years. What lets these countries work less and still produce more is what each hour has to work with: machinery, software, mature processes, decades of accumulated capital. The output comes from the system, not the stopwatch. They broke the link between hours and output, and took part of the winnings as time off. Where people work most and produce least At the other end, long hours are usually a sign of how much an economy still runs on people instead of capital. More than 1/2 of all jobs in Mexico and Colombia are informal, much of it self-employment in small operations where you put in long days for thin margins. The structural gap is enormous: capital per worker in Latin America is about 1/3 of the level in the United States and the OECD, and the region has lost ground, sliding from 46% of the OECD's output per worker in 1990 to 37% in 2023. This is not only a developing-economy story. Greece works the longest hours in the European Union and still lands near the floor of this chart, for the same structural reasons: lots of small firms, heavy self-employment, and labor-intensive work in tourism and agriculture. Here, long hours are a symptom of low productivity, not a cure for it. The United States breaks the pattern One big economy sits stubbornly above the line: the United States. It works long hours, about 1,805 a year, more than any wealthy Western European country, and still produces about $97 per hour, near the top of the table. That is the exact combination the rest of the chart says should not exist. Part of it is real. U.S. output per hour rides on heavy investment in information technology and intangible capital, which is why the U.S. number is not the accounting illusion Ireland's turns out to be. Part of it is just more time on the clock. The U.S. is the only advanced economy with no federally guaranteed paid vacation or holidays, while the EU mandates at least 4 weeks. The sharpest way to see it is the U.S. against Germany. A U.S. worker puts in about 474 more hours a year, nearly 12 extra 40-hour weeks, and at the end of all of it produces slightly less per hour than the German who clocked out earlier. The U.S. buys its high output with a blend of real productivity and a pile of extra hours nobody requires it to give back. The top of the chart deserves a second look The dots floating highest all come with an asterisk, for different reasons. Ireland tops the chart at $164 an hour, but that figure is mostly bookkeeping. In 2015 Irish GDP jumped 26% in a single year, not because the country got 1/4 more productive overnight but because multinationals shifted intellectual property onto Irish books. The distortion was blatant enough that Ireland's own statisticians built a separate measure, modified gross national income, to scrub it out. By that cleaner gauge, the real economy is something like 40% smaller than headline GDP suggests. Luxembourg, just below it at about $134 an hour, has a different quirk. Close to 1/2 of its workforce commutes in from France, Belgium, and Germany every day. Their output counts toward Luxembourg's GDP, but they are not part of its resident workforce, so output per worker comes out inflated. Norway, essentially tied with Luxembourg near the top, is the most honest of the three, and still worth a caveat. Its number is not an accounting illusion; the output is real. But a big share of it comes from oil and gas, Norway's largest sector by value added and exports. That output flows from offshore rigs run by relatively few people, which lifts the national figure without saying much about a typical Norwegian's workday. Strip out petroleum and mainland Norway, where most Norwegians actually work, looks a good deal more ordinary. The highest dots all need context: two are accounting, the third is oil. What this chart can and cannot say The tempting read is a life hack: work less, produce more. The chart does not say that, and dressing it up that way would be the dishonest version. The cleaner explanation runs the other direction. Rich, high-productivity countries can afford to work less, so as incomes climb people buy back their time through shorter weeks and longer vacations, while poorer economies put in long hours out of necessity. The causation runs both ways at once, and national income is doing most of the work behind the slope. There is a measurement quirk worth naming, since the y-axis has hours sitting in its denominator. GDP per hour is output divided by hours worked, so plotting it against hours carries a built-in downward tilt. That is not what drives the picture, though. The same negative pattern shows up against GDP per capita, which has no hours in it at all, and output per worker varies far more across these countries than their hours do. The tilt is real, not an accident of the math. One more honest flag: the OECD warns that its annual-hours figures are not built for comparing the exact level of hours between countries in a single year, partly because the measure folds in the self-employed, who tend to report long days. So some of the spread at the bottom of the chart comes from how the data is collected, not from people truly working more. This is not a productivity tip. It is a picture of how wealth, economic structure, and even the act of measurement all happen to point the same way. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It pulled average annual hours worked and GDP per hour worked for 40 economies directly from the OECD, matched each country to the latest year both series were available (2023 or 2024), fit the relationship, and labeled every country. The design iterated until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data sources: OECD average annual hours actually worked per worker and GDP per hour worked (current prices, current PPPs). The full dataset used for this chart is available here. --- ## U.S. World War II veterans are fading from living memory URL: https://www.randalolson.com/2026/06/01/wwii-veterans-fading/ Published: 2026-06-01 Updated: 2026-06-20 Categories: data visualization Tags: beautiful-charts-with-ai, world war ii, veterans, d-day, demographics Of the 16.4 million who served in World War II, the VA projects about 31,000 U.S. veterans are still alive in 2026, nearing zero by the early 2040s. Part of Teaching an AI Agent to Make Beautiful Charts Roughly 16.4 million people served in the U.S. armed forces during World War II. As of 2026, the Department of Veterans Affairs projects that about 31,000 of them are still alive. That is fewer than 1 in 500 of the people who wore the uniform, and the number is dropping fast. This week marks 82 years since the D-Day landings, which makes it a fitting moment to look at how quickly the last of that generation is leaving us. 16.4 million served. About 31,000 are left. Of those 16.4 million, more than 405,000 died during the war itself, according to the National WWII Museum. The survivors came home, built careers and families, and for decades were a visible part of civic life. By 2026, the VA projects only about 31,000 are still living. That is under 0.2% of everyone who served. In 2023 the same model counted about 95,000. The decline is accelerating This is not a gentle slope anymore. The VA's VetPop2023 model counted about 95,000 living WWII veterans in fiscal 2023 and projects about 31,000 by fiscal 2026. That is more than two-thirds of the 2023 survivors, gone in 3 years. The reason is straightforward: the youngest WWII veterans are now in their late 90s, and most are past 100. The VA's figures imply the count is falling by more than 14,000 a year right now, about 40 every day. D-Day's last witnesses On June 6, 1944, nearly 160,000 Allied troops came ashore in Normandy, and about 73,000 of them were from the United States, per the National WWII Museum. They were the visible, living link to the largest amphibious invasion in history. The 80th anniversary in 2024 was widely described as likely the last major D-Day commemoration with living veterans present. CNN reported that even the youngest survivors were nearing 100, and that organizers treated the gathering as a final chance to honor them in person. Two years later, that window has narrowed further. By the 2030s, a vanishing few The projection does not level off at a few thousand. It keeps falling. The VA expects fewer than 10,000 living WWII veterans by 2029 and fewer than 1,000 by 2034. At that point the survivors would be about 107 years old or older. A handful of supercentenarians may stretch the timeline slightly, but the trajectory is set by basic demography, not by any single uncertain assumption. When the last veteran dies, World War II leaves living memory By the early 2040s, the VA projects that essentially no living WWII veterans will remain. When the last one dies, World War II will pass entirely out of living memory, the way earlier wars already have. The last U.S. veteran of World War I, Frank Buckles, died in 2011 at age 110. The last surviving Union veteran of the Civil War, Albert Woolson, died in 1956. Each death marked the moment a war stopped being something anyone alive had seen firsthand. World War II is now approaching that same threshold within the next 20 years. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It researched the data, built the chart in Python, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: U.S. Department of Veterans Affairs, National Center for Veterans Analysis and Statistics, VetPop2023 (living veterans by period of service, Table 2L). The figures of 16.4 million who served and 405,399 wartime deaths come from the National WWII Museum. The dataset used for this chart is available here. --- ## How chess openings rose and fell, 1850 to 2026 URL: https://www.randalolson.com/2026/05/28/chess-openings-over-time/ Published: 2026-05-28 Categories: data visualization Tags: data visualization, chess, chess openings, chess history A streamgraph of 1.2 million master-level games shows the King's Pawn losing ground, the Queen's Pawn rising, and the Sicilian taking over after 1945. Back in 2014 I built a stacked area chart of chess opening popularity from 1850 to that year. It showed Black's reply to White's first move and stopped there: when I tried to push it to the next ply, the ribbons stacked into illegible noise. With another decade of master games on hand and a recursive streamgraph that keeps the same time axis at every depth, the version below picks up where that one stalled, and every ribbon opens into the lines beneath it. The 100-year d4 challenge In the 1860s, 93% of the games in this corpus opened with 1.e4. By the 1940s that share had fallen to about 42%, with 1.d4 right alongside it at roughly the same number. Getting there took about 80 years of theory, three generations of world champions, and a completely different way of thinking about the center of the board. The arc lines up with the eras chess writers have drawn for over a century. The Romantic era of Adolf Anderssen and Paul Morphy treated the center as ground to be seized in tactics; Morphy opened 1.e4 as White and answered 1.e4 with 1...e5 as Black, so nearly all of his games were open games. The Classical school of Wilhelm Steinitz, the first official world champion in 1886, and Siegbert Tarrasch traded sacrifices for positional principles but kept 1.e4 as the default. What finally broke the open-game monopoly was 1.d4, and a school that thought about the center in a completely different way. The hypermodern wave Click into 1.d4 in the chart and watch the 1...Nf6 ribbon. In the 1900s that reply turned up in under 3% of 1.d4 games. By the 1920s it was more than half of them, and it has been the majority answer to 1.d4 ever since. The Hypermodern school of the 1920s, named in Savielly Tartakower's 1924 book Die hypermoderne Schachpartie and codified by Aron Nimzowitsch's My System (1925–1927) and Richard Réti's Die neuen Ideen im Schachspiel (1922), made one claim above all: don't occupy the center with pawns, let your opponent push them out there and then attack them from the flanks with pieces. The movement crested in the 1930s. That decade is 1.d4's high-water mark in the whole chart: it reached about 52% of all master games while 1.e4 fell to 33%. The Indian Defenses are what the theory looks like over the board. Black answers 1.d4 with 1...Nf6 instead of 1...d5, then often plays ...g6 or ...e6, fianchettoes a bishop, and hits White's center from the wings. The Nimzo-Indian, the King's Indian, and the Grünfeld all live inside that ribbon. In barely one generation, the hypermoderns took a piece of board geometry the Romantics had never seriously contested and made it the most common reply to 1.d4 by a wide margin. The Sicilian takeover The most dramatic single line in this corpus is what happened to 1.e4 c5. The Sicilian Defense sat between 5 and 10% of 1.e4 games for the chart's first 70 years. Then it climbed, and it did not stop: by the 1960s it was over 40% of every game that began 1.e4, and it has hovered between roughly 40 and 50% ever since. In the same stretch the open game 1.e4 e5 (the Italian, the Ruy Lopez, and the rest) fell from about 85% of 1.e4 games in the 1860s to barely a fifth by the 1990s. The Sicilian's rise tracks a generation of players who built their names on it. Miguel Najdorf popularized the sharp 5...a6 line, which already existed, from the 1940s on, and it ended up carrying his name. Bobby Fischer made the Najdorf his main weapon against 1.e4 through the 1960s and early 1970s. Garry Kasparov leaned on the Sicilian, usually the Najdorf or a Scheveningen, as his main defense against Anatoly Karpov's 1.e4 across five world championship matches between 1984 and 1990. The open game was never refuted. The Sicilian just gave attacking players better practical results, and the field followed. The fall of 1.e4 e5 has a twist worth chasing in the chart. Click into it and run it forward: after bottoming out near a fifth of all 1.e4 games in the 1990s, it climbs back to almost 40% by the 2020s, as engines made rock-solid lines like the Berlin and the Petroff respectable again. Where things sit now From the 1970s on, the top of the streamgraph goes quiet. 1.e4 holds about 43% of all games, 1.d4 about 37%, 1.Nf3 (which often transposes into one of the other two) about 11%, and 1.c4 about 7%. Engines reshaped how openings get studied, from Deep Blue beating Kasparov in 1997 to AlphaZero in 2017, a network that learned the game only by playing itself and grew its own opening tastes, leaning toward 1.d4 and the English in its later training. But the top-level shares have drifted by single digits across the whole engine era, nothing like the 30-point swings of the first century. The action moved deeper. What changes now is which lines of the Sicilian and the Indian Defenses are in fashion, not whether to play them at all. Drill past the third or fourth move and the fights are everywhere: the rise and fall of the Najdorf English Attack, the Berlin Defense roaring back after Vladimir Kramnik used it to take Kasparov's title in 2000, the Ragozin slowly eating into the Queen's Gambit Declined. For the chart's first hundred years the question was which first move to play. For the last fifty it has been how deep the theory runs before someone finds new ground, and that is the question the recursive view was built to answer. Start at 1.e4 and see how far down you can go. How this chart was made The data is Lumbra's Gigabase OTB (CC BY-NC-SA 4.0), 1.2 million tournament games from 1850 to 2026 seen through a master-level strength lens: every recorded game before 1970, when FIDE Elo ratings began, plus games with both players rated 2400 or higher from 1970 on. The same strength lens was checked against the ChessGames.com data behind the 2014 post. The chart is built in D3.js as a recursive area chart: clicking a ribbon re-stacks its children on the same 177-year axis, the fix the 2014 version never had. --- ## The Misery Index Stopped Predicting the National Mood URL: https://www.randalolson.com/2026/05/27/consumer-sentiment-vs-misery-index/ Published: 2026-05-27 Categories: data visualization Tags: beautiful-charts-with-ai, consumer sentiment, misery index, inflation, vibecession By the misery index the U.S. economy is unremarkable, yet consumer sentiment is at an all-time low. That breaks a 40-year link no single number explains. Part of Teaching an AI Agent to Make Beautiful Charts Unemployment is near 4.3%. Inflation has cooled back under 4%. Add those together and you get the misery index, the back-of-the-envelope gauge of economic pain that has been around since the 1970s, sitting around 8%. That is an ordinary number. It is roughly where the misery index sat during the mid-2000s expansion. So why is the mood so dark? The University of Michigan's April 2026 final reading was 49.8, and the preliminary May 2026 number fell again to 48.2, the lowest reading on record. Lower than the depths of 2008. Lower than 1980, the worst the misery index has ever been. The economy looks unremarkable and the public feels worse than it did during the worst downturns on record. For 40 years, you did not see both at once. A 40-year relationship Economist Arthur Okun built the misery index in the 1970s as a quick read on how much the economy hurts: just the unemployment rate plus the inflation rate. Consumer sentiment is the other half of the picture, a monthly survey of how households say they feel about their money and the economy. The index is pinned to 100 at its strong-economy 1966 benchmark; its long-run average is around 85, and readings in the 50s show up only in crises like 1980 and 2008. From 1978 to 2021 sentiment and the misery index tracked each other. The relationship was loose, not a law (the misery index explains only about 40% of the month-to-month swings in sentiment, so the gray cloud is wide), but it always pulled the same direction. When the economy hurt, the mood sank with it, through the early-1980s inflation, the early-1990s recession, the 2008 crash. For 4 decades you could not be this gloomy without a painful economy to match. Then the dots fell off the line The red points are 2022 through 2026, with the slide starting in late 2021. They sit at a misery index near 8%, the same range as the calm years, but 30 to 40 points lower on sentiment than that range ever produced before. In plain statistical terms, recent months land 2 to 4 standard deviations below the historical fit and have stayed there for years. 1980 makes the size of it concrete. That May the misery index hit 22%, inflation running near 14% on top of a 7%-plus jobless rate, and sentiment bottomed near 52. Today the misery index is roughly 8% and sentiment is about the same. The public feels as grim as it did in the worst economy on record, with about a third of the economic pain. The 2008 financial crisis is the one prior moment the mood fell as far below the misery index as it has now, when the shock to confidence outran the rise in unemployment and inflation. But that lasted a few quarters and closed within a year. The post-2021 detachment has held for 4 straight years, and it is the first time the relationship itself, not just the level, stopped holding. This is not the mood overshooting a still-working relationship. The relationship stopped describing the data. It is not a partisan story The obvious suspicion is that this is just partisans trashing the economy when the other side holds the White House. The presidencies say otherwise. Trump's first term sits squarely on the historical line, including the 2020 pandemic spike, when unemployment surged and the mood fell exactly as the old relationship predicts. The break shows up under Biden and continues under Trump's second term: a Democratic administration and a Republican one, the same detachment. Partisanship does widen the spread. A Brookings analysis attributes roughly 3.6 points of the depressed reading to partisan bias, with Republicans about 2.5 times more reactive than Democrats. But widening the spread is not the same as moving the average. If this were partisan grievance, the mood would flip with the party in power. It hasn't. So what broke? Economists call this the "vibecession." The tidy guess is that people are reacting to high prices, and there is something to it, though not in the way the misery index measures. The misery index uses the inflation rate, the change in prices over the past year, which has cooled. Households are looking at the price level, which has not: prices are about 27% higher than they were at the start of 2020 and are not coming back down. Surveys back this up. Brookings, drawing on work by Stefanie Stantcheva, finds people report perceived inflation around 7.1% against an actual rate near 3.4%, because they are pricing the cumulative jump, not the latest year. That said, no single replacement number cleanly stands in for the misery index here, and a chart that claimed one would be overfitting a vibe. The honest summary is narrower and more interesting: the most durable rule of thumb in popular economics quietly stopped working, and the mood is now anchored to how prices feel rather than to the textbook gauge of pain. A Federal Reserve study of verified retail purchases sharpens the puzzle: 43% of people said they were doing worse than in 2019 while actually buying more. An honest caveat The survey itself changed. The University of Michigan moved from phone to web interviewing in 2024, and the new method reads lower, by the university's own estimate around 6.6 points and by an independent estimate about 8.9 points, enough that analysts call it a structural break. It is worth knowing, but it does not rescue the old relationship. The dots had already fallen off the line in 2022 and 2023, before the change, and even crediting the full adjustment leaves recent readings about 3 standard deviations below the historical fit. How this chart was made An AI agent built these charts end-to-end as part of the Beautiful Charts with AI series. It pulled the University of Michigan Index of Consumer Sentiment, the unemployment rate, and the Consumer Price Index from FRED, computed the misery index as unemployment plus trailing 12-month inflation, fit the historical relationship on 1978 to 2021 only, and produced 2 figures. Each design iterated until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data sources: University of Michigan Surveys of Consumers via FRED series UMCSENT; U.S. Bureau of Labor Statistics via UNRATE and CPIAUCSL. The full monthly dataset used for these charts is available here. --- ## The 2026 Indy 500 was the closest finish in race history URL: https://www.randalolson.com/2026/05/26/indy-500-closest-finish-2026/ Published: 2026-05-26 Categories: data visualization Tags: beautiful-charts-with-ai, motorsports, indianapolis-500, sports Felix Rosenqvist beat David Malukas by 0.0233 seconds at the 2026 Indianapolis 500, breaking the 1992 closest-finish record that had stood for 34 years. Part of Teaching an AI Agent to Make Beautiful Charts On Sunday, May 24, 2026, Felix Rosenqvist beat David Malukas to the finish line at the 110th Indianapolis 500 by 0.0233 seconds, the closest margin in the race's history. The record it broke had stood since 1992. A 7-lap shootout decided it Two late cautions reshaped the race. Rookie Caio Collet's fiery crash with 8 laps to go brought a red flag, and a second caution on lap 197 after Mick Schumacher hit the SAFER barrier set up a one-lap shootout. Rosenqvist's Meyer Shank Racing teammate Marcus Armstrong took the white flag in the lead, with Malukas second and Rosenqvist third. Malukas powered into the Turn 1 lead on the final lap. Rosenqvist passed Armstrong on the high line out of Turn 4, then caught Malukas exiting the corner and won the drag race down the front stretch. The margin at the line was 0.0233 seconds, or roughly half an Indy-car length at 220 miles per hour. "It kind of worked out the right way when I got back to third, and then I just had to run a flat-out lap on the high line, and it stuck," Rosenqvist said afterward. It was his first Indy 500 win, and Meyer Shank Racing's first IndyCar victory since 2021. The 34-year-old record it finally beat The previous closest finish came from a brutally cold race day. On May 24, 1992, with a high of 58 degrees, Al Unser Jr. held off Scott Goodyear by 0.043 seconds on cold, slick tires that turned the day into a crash-filled marathon. Michael Andretti had led by nearly 30 seconds before a fuel pump failure ended his race 11 laps from the finish. Unser Jr. became the first second-generation Indy 500 winner, and the side-by-side finish-line photo became one of the iconic images of IndyCar racing. That margin held the record for 34 years. Rosenqvist's 2026 finish was 1.85 times closer. Every one of the 10 closest finishes has come since 1982 The top 10 list has no race from the first 7 decades of the Indy 500 on it. The oldest entry is 1982, when Gordon Johncock held off Rick Mears by 0.16 seconds. 6 of the top 10 have come since 2014. The modern Indy 500 produces sub-half-second finishes at a faster pace than any prior generation of the race. From 13-minute margins to thousandths of a second In 1913, French driver Jules Goux won the 3rd-ever Indianapolis 500 by 13 minutes and 8 seconds, drinking champagne at his pit stops to stay hydrated through the 6-hour-plus race. Through the 1920s and 1930s, margins were typically measured in minutes. Cars broke. Drivers had to slow for the simple act of finishing. By the 1960s, durability had improved and the typical margin had fallen to under a minute. By the early 1980s, sub-10-second finishes were common. Electronic transponder timing arrived at Indianapolis in 1990 and made decimal-place margins meaningful for the first time. What's done the rest is parity: common chassis, regulated aero, fuel-and-tire strategies that converge late in a race, and a depth of competition at the top of IndyCar that makes one-lap shootouts a recurring feature rather than a freak event. The chart spans more than 4 orders of magnitude: the 1913 margin of 788 seconds is roughly 34,000 times the 2026 margin of 0.0233 seconds. Track time, telemetry, and a more competitive field have collapsed the distance between winners and runners-up to the resolution of the timing system itself. How this chart was made An AI agent built these charts end-to-end as part of the Beautiful Charts with AI series. It compiled the Indianapolis Motor Speedway's official margin-of-victory data, added the 2026 result from IndyCar's race report, and iterated on the design until both charts passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: Indianapolis Motor Speedway's historical margin-of-victory page for every race with a recorded time, with the 2026 figure confirmed by IndyCar and ESPN. Years where the official margin is recorded in laps rather than seconds are excluded from the long-history chart. The CSVs used for these charts are available here: top 10 closest finishes and all margins 1911 to 2026. --- ## NOAA's hurricane outlook misses one way: it underforecasts big seasons URL: https://www.randalolson.com/2026/05/19/noaa-hurricane-outlook-vs-reality/ Published: 2026-05-19 Updated: 2026-06-20 Categories: data visualization Tags: beautiful-charts-with-ai, hurricanes, noaa, weather forecasting, climate From 2003 to 2025 the Atlantic named-storm total fell outside NOAA's May range 9 times, 8 of them underforecasts. NOAA's 2026 outlook lands May 21. Part of Teaching an AI Agent to Make Beautiful Charts On Thursday, May 21, NOAA will announce how many storms it expects in the 2026 Atlantic hurricane season, 11 days before the season officially starts on June 1. Since 2003, the actual named-storm count has landed inside NOAA's May predicted range about 6 years in 10. When it missed, it almost always missed the same way: the season was busier than NOAA called for. The forecast misses in one direction Across the 23 seasons from 2003 through 2025, the actual total landed inside NOAA's May named-storm range 14 times. Of the 9 years it missed, 8 were underforecasts: the season produced more storms than the top of NOAA's range. Only 2006 went the other way. NOAA states that activity should fall within each predicted range in about 70% of seasons with similar conditions. For named storms, the observed hit rate is closer to 61%, and the shortfall is almost entirely seasons that turned out busier than predicted. When NOAA's May outlook is wrong, the season is almost always worse than forecast, not quieter. The error has a consistent sign: the years NOAA missed were the hyperactive ones, not the quiet ones. 2005: forecast 12 to 15, then 28 storms NOAA's May 2005 outlook called for 12 to 15 named storms and a 70% chance of an above-normal season. The season delivered 28, the most on record at the time, including Katrina, Rita, and Wilma. NOAA's own post-season assessment pinned the explosion on tropical-Atlantic sea-surface temperatures reaching an 1870 to 2005 record, La Nina-like patterns in the Pacific, and exceptionally weak wind shear, all stacked on top of the active era that began in 1995. By early August 2005, NOAA had revised its outlook up to 18 to 21 named storms. The season still came in nearly double the top of the May range. 2020: the record NOAA didn't see in May The May 2020 outlook predicted 13 to 19 named storms with a 60% chance of an above-normal season. The Atlantic produced 30, a new record. The biggest driver was the warmest Atlantic sea-surface temperatures on record, compounded by a developing La Nina that cut wind shear and an enhanced West African monsoon. In August, NOAA escalated to an "extremely active" call of 19 to 25 named storms, the first time it had ever forecast as many as 25. Even the dramatically raised August ceiling still came in 5 storms short of the 30 that formed. 2006: the only year the forecast was too high 2006 is the lone dot below its forecast bar. NOAA's May 2006 outlook called for 13 to 16 named storms and an 80% chance of an above-normal season. Only 10 formed. A moderate El Nino developed rapidly and unexpectedly in late summer, ramping up wind shear, while a dry, dust-laden Saharan Air Layer suppressed activity through August, per Tropical Storm Risk's season verification and the season summary. The single overforecast in 23 years came from an El Nino that arrived faster than the May outlook assumed. That is the same lever, ENSO, that drives the underforecasts, just pulled in the opposite direction. Why the misses run the same way The structural reason sits in the calendar. NOAA issues the outlook in May, but spring is the worst time of year to predict El Nino and La Nina: forecast models hit what NOAA's Climate Prediction Center calls the spring predictability barrier, explaining less than a third of ENSO variability from an April vantage point. NOAA's own verification work shows the May outlook has little statistical skill, while the August update is far more accurate once the season's drivers are locked in. The bias toward underforecasting is the active era showing through. Since 1995 the Atlantic has been in a high-activity phase, so when a forecaster anchors near a moderate May midpoint and conditions turn favorable, there is far more room to blow past the top of the range than to fall below the bottom of it. The downside risk on these forecasts is one-directional. What the 2026 outlook is up against Colorado State University's April 2026 forecast calls for a somewhat below-normal season: 13 named storms, 6 hurricanes, and 2 major hurricanes, with overall activity around 75% of the long-term average. The reasoning is a weak La Nina expected to flip to a moderate or strong El Nino, which raises Atlantic wind shear, paired with a mixed sea-surface-temperature pattern. NOAA's official 2026 numbers arrive May 21. A below-normal call has historically been the safer kind to make, but 2005 and 2020 show the rare big miss has almost always run the other way: a season far worse than the May outlook, not quieter. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It compiled NOAA's May preseason forecast ranges and the National Hurricane Center's end-of-season totals for 2003 through 2025, classified each year as within, above, or below the forecast range, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: NOAA Climate Prediction Center May Atlantic hurricane outlooks for the forecast ranges and NOAA National Hurricane Center season totals for the actuals. This analysis covers 2003 through 2025, the period over which NOAA's May outlook consistently stated an explicit predicted range for the upcoming season; earlier outlooks (2000 to 2002) framed activity in probabilistic terms and are not included. The CSV used for this chart is available here. --- ## Inflation worry flipped party lines while inflation itself barely moved URL: https://www.randalolson.com/2026/05/14/inflation-worry-flipped-parties/ Published: 2026-05-14 Categories: data visualization Tags: beautiful-charts-with-ai, inflation, public opinion, pew research, cpi, partisanship Pew Research finds Republicans and Democrats traded places on inflation worry from 2024 to 2026, even though CPI rose less than a percentage point. Part of Teaching an AI Agent to Make Beautiful Charts In May 2024, 6 months before the presidential election, 80% of Republicans and 46% of Democrats told Pew Research that inflation was a "very big problem" facing the country. In April 2026, 15 months into Trump's second term, the 2 numbers are 55% and 74%. The parties traded places. Headline inflation in that window moved from 3.3% to 3.8%. In 2022, both parties' worry rose with inflation When Pew first asked the question in May 2022, U.S. headline CPI was at 8.6% year over year and climbing toward a June 2022 peak of 9.1%. Republicans hit 84% and Democrats hit 57%, both their highest readings of the series. The 27-point partisan gap was already there. This is the only wave where both parties' worry and the actual inflation rate move in the same direction. In every subsequent wave, partisan worry and inflation pull apart. By 2024, the gap widened as inflation cooled From May 2022 to May 2024, CPI fell from 8.6% to 3.3%, a 5.3 point drop that brought inflation closer to the Fed's 2% target. Republican worry fell 4 points, from 84% to 80%. Democratic worry fell 11 points, from 57% to 46%. The partisan gap widened from 27 points to 34 points as the actual problem shrank. The lines crossed within a year of Trump taking office The next wave landed in February 2025, 2 weeks after the inauguration. Republican worry was at 73%, down 7 points from May 2024. Democratic worry was at 53%, up 7 points from May 2024. By April 2026 the crossover was complete: 55% of Republicans and 74% of Democrats called inflation a very big problem. From May 2024 to April 2026, Republican worry fell 25 points and Democratic worry rose 28 points. That is a 53 point swing in 2 years. Inflation did rise during the Trump months. CPI climbed from 2.8% in February 2025 to 3.8% in April 2026, the highest reading since May 2023. Gas prices jumped 50% after the war with Iran started, and tariffs on Mexican produce pushed tomato prices up roughly 40% year over year. Over the same window as the 53-point partisan swing, headline CPI rose half a percentage point. Inflation itself barely moved The bottom panel tells the part of the story the perception lines do not. Between the May 2024 Pew wave and the April 2026 wave, headline CPI moved half a percentage point. CPI went from 8.6% down to 3.3% during Biden's term, then back up to 3.8% in Trump's first 15 months. The Trump-era inflation level looks a lot like the late-Biden inflation level. What changed is who is in office. The number did not. This is what partisan motivated reasoning looks like Political scientists have been documenting this pattern for decades. Brady, Ferejohn, and Parker's 2022 paper finds that partisan bias in U.S. economic perceptions has grown substantially over the past 20 years, with supporters of the in-party rating the economy far more favorably than supporters of the out-party even after controlling for actual conditions. The most recent precedent is from 2017. Gallup's Economic Confidence Index showed Republicans jumping 77 points and Democrats dropping 38 points in the year after Trump's first inauguration, with the underlying economy barely changed. Pew's inflation question is the same phenomenon on a different topic and a different year. When the party in the White House changes, the perception of the economy changes with it, even when the data does not. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It pulled the Pew Research National Problems waves from May 2022 through April 2026, computed headline CPI year-over-year changes from the BLS CPI-U series, rendered the 2 metrics as stacked time-aligned panels, and iterated on the layout until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data sources: Pew Research Center, "Americans See Health Care Costs, Inflation as Big Problems" (American Trends Panel Wave 192, fielded April 20-26, 2026, n=5,103); FRED CPIAUCNS, the U.S. Bureau of Labor Statistics Consumer Price Index for All Urban Consumers, not seasonally adjusted. The CSV used for this chart is available here. --- ## U.S. measles cases broke the post-elimination floor in 2025 and 2026 URL: https://www.randalolson.com/2026/05/13/us-measles-elimination-status-at-risk/ Published: 2026-05-13 Updated: 2026-06-20 Categories: data visualization Tags: beautiful-charts-with-ai, measles, public health, vaccination, cdc, mmr The U.S. reported 2,288 measles cases in 2025 and 1,842 by May 7, 2026, 23x the typical post-elimination year. WHO reviews elimination status in November. Part of Teaching an AI Agent to Make Beautiful Charts In 2000, the U.S. did something hard: it stopped measles. Endemic transmission halted, and for the next 24 years the country averaged 79 confirmed cases annually, almost all of them imported by travelers and stamped out before they spread. That run is over. 2025 brought 2,288 cases, the worst year since 1991. 2026 is already at 1,842 with 2/3 of the year still ahead. The elimination status the U.S. earned a generation ago is up for WHO review this November. The post-elimination floor barely moved for 24 years From 2001 through 2024, the median U.S. year saw 79 measles cases, and 19 of those 24 years came in under 200. The CDC and PAHO declared measles eliminated in 2000 because the country had stopped having sustained domestic transmission chains. WHO defines elimination as the absence of any chain that lasts 12 months or longer. That low post-2000 baseline is what elimination looks like in practice: sporadic imported cases that fizzle before they spread. 2014 and 2019 were shocks that didn't break the floor The 2014 and 2019 outbreaks pushed cases above the baseline but never threatened elimination status. The 2014-15 Disneyland outbreak infected 147 people across 7 states. 45% of California patients were unvaccinated, and 67% of the vaccine-eligible patients were intentionally unvaccinated. California responded by passing SB 277 in 2015, which ended personal-belief exemptions for school enrollment. The 2019 New York and New Jersey outbreak was larger and more concentrated. 1,274 cases nationally, with 93% of New York City cases in the Orthodox Jewish community and 73% of patients residing in Williamsburg, Brooklyn. The median patient age was 3. NYC Mayor de Blasio declared a public health emergency on April 9, 2019, and New York repealed its religious vaccination exemption in June. Both years cleared the chain before the 12-month elimination clock ran out. 2025 and 2026 are the West Texas chain The 2025 surge started in Gaines County, Texas, a Mennonite community where roughly 20% of kindergartners claim a vaccine exemption compared to the state average of less than 4%, per Texas Tribune reporting. By March 2025 the outbreak had spread across the Southwest, eventually killing 2 unvaccinated children in Texas and 1 unvaccinated adult in New Mexico, the first U.S. measles deaths in a decade. The CDC's May 7, 2026 update counts 1,842 cases across 39 jurisdictions, with 92% of patients unvaccinated or of unknown vaccination status. The Texas-linked transmission that started in January 2025 is what puts elimination status in question. WHO defines elimination as the absence of a transmission chain lasting 12 months or longer, and the Region of the Americas measles commission will assess the U.S. at its next annual meeting. National MMR coverage has dropped below the herd immunity threshold The West Texas chain caught hold because national vaccination coverage is no longer enough to prevent local outbreaks from spreading. CDC's SchoolVaxView shows kindergarten MMR coverage fell to 92.5% in the 2024-25 school year, down from 95% pre-pandemic and below the Healthy People 2030 target of 95% needed to prevent community transmission. 39 states now sit below the 95% threshold, up from 28 before the pandemic. Vaccine exemptions, almost all non-medical, rose to a record 3.6% of kindergartners in 2024-25. That works out to roughly 286,000 unvaccinated kindergartners nationally. Pockets of much lower coverage drive most of the risk: Idaho's statewide kindergarten MMR rate is 78.5%, and individual counties run lower than that. What losing elimination status would actually mean The Region of the Americas measles commission already lost its regional elimination status in November 2025, after Canada's outbreak chain crossed 12 continuous months. The U.S. retained its country-level status at that meeting and faces its next assessment in November 2026. If the Texas-linked chain that started in January 2025 has not been definitively interrupted, the U.S. will join Canada. Losing the designation does not change daily clinical guidance. MMR is still the recommended vaccine, exposures are still reported to the CDC, and outbreak response still runs through local health departments. What changes is the global label. Elimination is recoverable. The United Kingdom lost its status in 2019, regained it in 2021 when COVID restrictions suppressed transmission, and lost it again in early 2026 after cases rebounded through 2024. Brazil lost elimination in 2019 and earned it back in 2024 after 5 years of sustained vaccination work. The November review will just label what the chart already shows. What changes next year's chart is whether kindergarten MMR coverage climbs back above 95%. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It pulled CDC annual measles case counts via Our World in Data, computed the 2001-2024 post-elimination median, rendered the result as a bar chart with that empirical baseline drawn directly on the plot, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: U.S. Centers for Disease Control and Prevention measles case counts, mirrored by Our World in Data. The 2026 figure is year-to-date through May 7 per the CDC's latest update. The CSV used for this chart is available here. --- ## 1 in 5 U.S. registered nurses are now 65 or older URL: https://www.randalolson.com/2026/05/12/us-rn-workforce-aging/ Published: 2026-05-12 Updated: 2026-06-20 Categories: data visualization Tags: beautiful-charts-with-ai, nursing workforce, ncsbn, registered nurses, healthcare workforce, hrsa The U.S. registered nurse workforce has split into two extremes: 1 in 5 RNs are now 65 or older, and the under-30 share fell to a 10-year low. Part of Teaching an AI Agent to Make Beautiful Charts The median U.S. registered nurse is now 50 years old. The 2024 NCSBN National Nursing Workforce Survey, released in April 2025 and based on responses from 744,714 RNs, shows the age distribution has pulled apart at the ends. 1 in 5 RNs are now 65 or older, and the share of RNs under 30 fell to a 10-year low. The chart below pairs the 2015 share of each 5-year age cohort with the 2024 share. The hollow dot marks 2015, the filled dot marks 2024. The middle bands barely moved. The two extremes did almost all the work. 1 in 5 RNs are now 65 or older The 65+ share of the RN workforce grew from 12.4% in 2015 to 18.3% in 2024, a 5.9% jump and the single largest move in the distribution. The path was anything but straight. The 65+ share spiked to 19.0% in the 2020 survey as older nurses delayed retirement through the early pandemic, fell to 13.2% in 2022 as exhausted veterans left, then rebounded to 18.3% in 2024 as some returned. NCSBN frames the recent rise as workforce stabilization rather than continued aging. The same survey warns that "many of the last of the baby boomers are expected to exit the workforce in 2027 as they reach retirement age." 39.9% of RNs say they intend to leave the workforce or retire within 5 years, and 41.5% of those name stress and burnout as the cause. Most of the 65+ cohort visible in the chart is one retirement decision away from exiting. Why the under-30 share isn't bouncing back At 7.9%, the under-30 share of the RN workforce hit its lowest level in the 10-year NCSBN series, below the prior low of 8.4% set in the 2020 survey. It would be easy to read this as a drop in interest from young people. Application data says otherwise. The American Association of Colleges of Nursing reports that U.S. nursing schools turned away 80,162 qualified applications from baccalaureate and graduate programs in 2024. BSN enrollment grew 4.9% that same year, adding more than 12,000 new students. The bottleneck is faculty shortages, clinical placement sites, classroom space, and budget. About 1/3 of nursing faculty is expected to retire by 2025, which tightens the constraint right when the workforce most needs to expand it. The 55-59 dip is the boomer wave moving through The 55-59 row is the cleanest demographic mechanism in the chart. The 55-59 cohort fell 3.6 points, from 13.6% in 2015 to 10.0% in 2024. The same population that was 55-59 in 2015 is now 64-68, and the bulk of it has aged into the 65+ band. The size of the 55-59 drop is a near-mirror of the 65+ rise, which is what you would expect when a large cohort moves through a fixed distribution and few people enter behind it. What happens after 2027 HRSA's 2024 Nurse Workforce Projections estimate an 8% shortage of registered nurses by 2028, narrowing to 3% by 2038 as the pipeline catches up. The national average hides the geography. By 2038, rural areas face an 11% RN shortage while metropolitan areas face 2%. More than 1 million nurses are projected to retire by 2030, which is most of the 65+ cohort shown at the bottom of the chart. The pipeline is responding. Newly licensed RNs grew 36.1% from 2019 to 2023. Until nursing schools can seat the qualified applicants they currently turn away, the under-30 share is unlikely to recover and the 65+ share will keep doing the heavy lifting. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It pulled Table 6 of the NCSBN 2024 National Nursing Workforce Survey, computed the share for each 5-year age cohort in 2015 and 2024, rendered the pair as a dumbbell, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: Table 6 of "The 2024 National Nursing Workforce Survey", Journal of Nursing Regulation, April 2025 supplement, from the National Council of State Boards of Nursing. HRSA projection figures come from the 2023-2038 Nurse Workforce Projections fact sheet and the 2024 State of the Health Workforce Report, both from the U.S. Bureau of Health Workforce. Nursing school capacity figures come from the American Association of Colleges of Nursing Nursing Faculty Shortage fact sheet. The CSV used for this chart is available here. --- ## U.S. federal workforce hits 60-year low after fastest peacetime cut URL: https://www.randalolson.com/2026/05/09/us-federal-workforce-fastest-peacetime-cut/ Published: 2026-05-09 Categories: data visualization Tags: beautiful-charts-with-ai, federal employment, bureau of labor statistics, federal workforce, civil service, bls Year over year, U.S. federal civilian employment is shrinking faster than at any point since 1939 outside wars and the decennial Census. Part of Teaching an AI Agent to Make Beautiful Charts Federal civilian employment has only fallen this fast a handful of times since the Bureau of Labor Statistics started counting in 1939. Every prior episode has a tidy explanation. WWII demobilization in 1946. Korean War demobilization in 1953-54. The pulse of temporary Census workers coming off the rolls after each decennial Census. The contraction running through April 2026 is the first that doesn't fit any of those buckets. The April number, released by BLS on May 8, put the federal civilian payroll at 2,665,000 workers. That is 348,000 below the October 2024 peak and the smallest the federal civil service has been since May 1966. The cut is also the deepest peacetime, non-Census 12-month drop in the entire BLS series. Year over year, the federal workforce shrank 10.5%. The previous post-1950 record holder, the Clinton-Gore reinvention of the late 1990s, peaked at 3.4%. The 1980s and 1990s mirage The chart looks like there are 2 clear precedents for the current cut. Around 1981 the line plunges almost as deep. So does 1991. If you've ever read a sentence like "the federal workforce shrank in the early 1980s" or "after the Cold War", these are the troughs the writer was looking at. Neither of them is what most readers think it is. Both are dominated by the wind-down of temporary Census workers that the federal government had hired in the spring of 1980 and the spring of 1990. April 1980 alone added roughly 300,000 temporary workers to the federal headcount; when those workers came off the rolls a year later, the year-over-year change went deeply negative whether or not the standing civil service had moved at all. The Reagan-era reductions in force were real, but the standing-workforce drop from his January 1981 inauguration to the 1981-82 recession trough in May 1982 was 106,000 workers, or 3.6%. The PATCO firings of August 1981 accounted for about 11,000 of that. The rest was attrition and a hiring freeze. Strip out the Census, and the deepest peacetime year-over-year drop on record before 2025 was 3.4%. The current trough at 10.8% is just over 3x that. What 18 months looks like next to 6 years Federal employment is sticky. Civil servants have job protections, agencies have appropriations cycles, and most administrations move the standing workforce by a few percent over a presidential term. The standard comparison for a deliberate peacetime cut is the Clinton-Gore reinvention. Vice President Gore launched the National Performance Review in March 1993 with a target of 252,000 federal positions cut over 5 years. By January 1999 the federal civil service had shrunk by 327,000 workers, a reduction of 10.6% over 6 years. The current contraction has cut 348,000 workers in 18 months. Same magnitude, a quarter of the time. A deferred-resignation offer that ran for 2 weeks pulled about 154,000 people off the standing federal payroll in a single batch. The Fork in the Road On January 28, 2025, the new administration sent a memo to most of the federal civil service titled "Fork in the Road." It offered a deferred resignation: leave by September 30 and continue drawing full pay and benefits in the meantime. Acceptance had to be filed within 2 weeks, by 7:20 PM Eastern on February 12. About 154,000 employees took the deal and stayed on administrative leave for 6 months or longer; administration projections put the total higher than 200,000 once later waves were counted. That memo did most of the work. The rest came from formal Reductions in Force, which hit hardest at agencies the administration had been most public about wanting to shrink. The Department of Health and Human Services announced 20,000 cuts in March 2025, about a quarter of the department. Smaller RIFs hit the Department of Education, the Consumer Financial Protection Bureau, USAID, and others, sometimes followed by court orders that paused or reversed them. A government-wide hiring freeze let attrition do the rest of the work without making news. The cuts now look to be slowing. The April 2026 monthly drop was 9,000 workers, against monthly drops of 25,000 to 50,000 through the spring and summer of 2025. The 12-month rolling change has eased from 10.8% in January to 10.5% in April. Some agencies are quietly rehiring people they cut, after discovering they could not perform basic functions without them. The fastest peacetime cut on record may be approaching its floor. Why the Census shows up so often in this chart One detail in the chart is worth a closer look: the deepest non-WWII drops in the entire 87-year series are both Census artifacts. The 2010 Census peaked at 564,000 temporary workers in May 2010, the 2020 Census peaked at 288,000 in August 2020 (smaller because of online self-response and pandemic compression), and every prior cycle had a comparable spike. Per BLS analysis, those temporary workers come on the federal payroll in the spring of years ending in zero and come off it a year later. The rolling year-over-year change picks up the difference and registers it as a contraction even though the standing civil service has been steady the entire time. This matters for any historical comparison of federal workforce trends. The 1981 and 1991 troughs in the chart are mostly the 1980 and 1990 Census wind-downs, not policy choices. The current contraction has no Census interaction at all. Its 12-month window from April 2025 to April 2026 contains zero temporary-Census-worker hiring or release activity. The standing civil service did all the moving. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It pulled the BLS Current Employment Statistics series for federal civilian employment from FRED, computed the year-over-year percentage change for every month from 1939 through April 2026, classified each trough by historical episode, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data sources: FRED series CES9091000001 from the U.S. Bureau of Labor Statistics Current Employment Statistics, "All Employees: Federal", seasonally adjusted, monthly, January 1939 through April 2026. The April 2026 figure comes from the May 8, 2026 Employment Situation release. The CSV used for this chart is available here. --- ## World Happiness Report 2026: what income can't explain URL: https://www.randalolson.com/2026/05/07/whr-2026-happiness-beyond-income/ Published: 2026-05-07 Updated: 2026-06-20 Categories: data visualization Tags: beautiful-charts-with-ai, happiness, world happiness report, central america, east asia, income inequality In the 2026 World Happiness Report, all but one Latin American country scores above what income predicts. East Asia shows the opposite pattern. Part of Teaching an AI Agent to Make Beautiful Charts Costa Rica ranked #4 in the 2026 World Happiness Report, its highest finish ever. Its GDP per capita is about $27,000 in purchasing-power terms, less than a quarter of what Singapore earns and about a third of the U.S. figure. Yet Costa Rica sits above Sweden, above Australia, above Germany. Strip out what income explains, and the real story is even stranger. Latin America's 19-for-20 Run a regression of happiness on income across all 138 countries with data and you get a fairly predictable line: richer countries are happier. Every Central American country in the dataset sits well above that line. So does nearly every other country in Latin America. Of the 20 Latin American countries with World Bank income data, 19 score higher than income alone would predict. Trinidad and Tobago is the lone exception, barely below the prediction at -0.29 points. This pattern has held for over a decade. Mariano Rojas, writing in the 2018 World Happiness Report, found that Latin American life evaluations run above what the standard six-variable model predicts, a gap later quantified at roughly 0.6 points. His analysis identifies the quality and abundance of family ties as the main driver. In Latin America, family functions as the primary social safety net rather than state institutions, and the data shows it counts for a lot in self-reported happiness. That shows up at the extreme end too. Nicaragua has a GDP per capita of about $7,700, low enough that a strict income-to-happiness model would predict a score around 5.1. Its actual score is 6.3. A study of 99 garbage collectors in Leon, Nicaragua found zero correlation between monthly income (ranging from under $25 to over $65) and happiness, while satisfaction with family and friends was by far the strongest predictor of wellbeing. The Gallup Global Emotions Report found that in 2024, Latin American countries dominated the top of the Positive Experience Index, with Paraguay and Panama leading at 86 out of 100 against a world average of 71. The East Asian shortfall The other end of the chart is just as striking. Japan, South Korea, Singapore, and Hong Kong all score well below what income predicts. Singapore has a GDP per capita of roughly $133,000, the highest in the dataset, yet its happiness score (6.59) sits about 0.70 points below the regression line. Hong Kong earns around $66,000 per capita but scores 1.18 points below prediction. Four of the five East Asian economies with income data fall below zero. China is the exception, coming in marginally above the line at +0.11 points. This pattern has a name. Yew-Kwang Ng identified it in 2002 as the "East Asian happiness gap", arguing that Confucian-influenced achievement orientation makes people drive for the next goal rather than enjoy the current one. A more recent review by Choi and Choi (2025) identifies five cultural mechanisms behind the gap. Among them: East Asians are more likely to evaluate their lives against external standards rather than internal feelings, and more sensitive to social comparison. Choi and Choi argue these are real wellbeing differences, not artifacts of self-reporting. Hong Kong adds structural factors on top of that cultural baseline. Its Gini coefficient is 0.539, among the highest of any developed economy. Housing prices run more than 14 times median household income, against a threshold of 3.0 for what demographers consider affordable. The political unrest of 2019 pushed depression prevalence from roughly 2% to 11% in a matter of months. What the gap tells us about money The regression line captures a real relationship. Richer countries do tend to be happier, and the correlation holds across the full 138-country dataset. But the residuals show how much variance income leaves unexplained. Costa Rica, earning about $27,000 per person, scores as high as the model predicts for a country earning $160,000. Nicaragua at $7,700 scores 1.2 points above what income would predict, one of the largest positive residuals in the dataset. Income predicts the floor, not the ceiling. The Latin American surplus and the East Asian shortfall have held across multiple WHR editions. GDP growth can shift a country's predicted happiness score, but it does not change which side of the regression line a country falls on. That part comes from somewhere else. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It pulled happiness scores from the World Happiness Report 2026, fetched GDP per capita PPP data from the World Bank API, ran an OLS regression of happiness on log(GDP), and computed residuals for 138 countries. The chart design iterated until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Note on coverage: Venezuela is excluded because the World Bank does not publish a current GDP figure for it. Taiwan is also excluded for the same reason. Afghanistan has a happiness score in the WHR (1.45) but no matching World Bank GDP entry and is excluded from the residuals analysis. Data sources: World Happiness Report 2026 (Cantril ladder scores, 3-year average 2023-2025); World Bank WDI indicator NY.GDP.PCAP.PP.KD (GDP per capita PPP, constant 2021 international dollars, 2024). The full dataset used for this chart is available here. --- ## Mexico ships 83% of every fresh avocado the U.S. imports URL: https://www.randalolson.com/2026/05/05/us-avocado-imports-mexico-cinco-de-mayo/ Published: 2026-05-05 Updated: 2026-06-20 Categories: data visualization Tags: beautiful-charts-with-ai, avocados, agricultural trade, mexico, cinco de mayo, supply chain In 2025, the U.S. imported 2.7 billion pounds of fresh avocados. Mexico supplied 83%. Peru, Colombia, and the Dominican Republic split the rest. Part of Teaching an AI Agent to Make Beautiful Charts Cinco de Mayo is the second-biggest avocado day in the U.S., behind only Super Bowl Sunday. About 81 million pounds will get smashed into guac, sliced onto tacos, or piled on toast today. Almost none of those avocados were grown here. The vast majority came from a single country, and even more specifically, from a single state inside that country. One country, one variety, near-total dominance The U.S. imported 2.7 billion pounds of fresh avocados in calendar year 2025, on top of about 430 million pounds grown at home (almost all of that in California). Of those imports, Mexico shipped 2.24 billion pounds, or 83% of the total. The next 4 suppliers combined did not reach a fifth of Mexico's volume. Mexico ships 11 times as many avocados to the U.S. as the next-biggest source, Peru. Almost every one of those Mexican avocados is the same variety: Hass. Hass traces to a single seedling that mail carrier Rudolph Hass planted in his La Habra Heights, California yard in 1926 and patented in 1935. A 2019 genetic study found the variety is roughly 61% Mexican and 39% Guatemalan ancestry. It now makes up about 95% of all U.S. avocado consumption. It wasn't always this way For most of the 20th century, no Mexican avocado could legally enter the country. The U.S. imposed a ban in 1914 over concerns about weevils, scabs, and other pests that could damage California orchards, and that ban held for almost 8 decades. It began to crack in 1993 with limited shipments to Alaska, then opened in stages through the rest of the 1990s and 2000s. Mexican Hass only got year-round access to all 50 states in 2007. Before that, Chile was the giant. Through most of the 1990s, Chilean avocados accounted for more than 80% of U.S. imports. Once Mexico finally walked in the front door, the picture flipped almost overnight. Chile's share fell from over 80% in the 1990s to less than 1% by 2025. The U.S. avocado market we have today is barely 30 years old. Why Michoacán wins Mexican avocado exports to the U.S. are not a national effort. They come almost entirely from one state, Michoacán, with a smaller contribution from neighboring Jalisco. The USDA's March 2026 Mexico Avocado Annual reports that Michoacán alone produced 75% of Mexico's 2.73 million metric tons in 2025. The state sits in the Trans-Mexican Volcanic Belt, where well-draining volcanic soil covers about 80% of the planted area. The geography gives Michoacán 4 overlapping blooming seasons in a single calendar year, which means somewhere in the state, an avocado is being harvested every week. That continuous supply is why U.S. grocery stores can keep Hass on the shelf 52 weeks a year with no visible seasonal gap. No other producing region in the world matches that combination of soil, elevation, climate, and bloom cycle at scale. The risk of leaning so hard on one source The flip side of that efficiency is fragility. In June 2024, the USDA paused every avocado and mango inspection in Michoacán for 10 days after 2 of its inspectors were assaulted and briefly held at a roadblock. A similar pause hit shipments in February 2022, days before the Super Bowl. Avocado growers in Michoacán have long reported extortion demands from cartels, sometimes thousands of dollars per acre, with the lucrative export trade as the obvious target. Every disruption pushes wholesale prices up almost immediately, because no other supplier can scale fast enough to plug the gap. You can see the market slowly hedging in the year-over-year numbers: Mexico held 89% of U.S. import volume in 2023, dropped to 87.6% in 2024, and sat at 83% by the end of 2025. Most of that lost share went to Peru and Colombia. Where the other 17% comes from Peru is the seasonal complement. Its harvest runs March through August, exactly when Mexican supply hits its annual low. About 86% of Peru's export crop sits along its dry Pacific coast under intensive drip irrigation, a fundamentally different growing environment than Michoacán's volcanic highlands. Colombia is the fastest-growing supplier, jumping from a rounding error a decade ago to 4.6% of U.S. imports in 2025. Its biggest gains land in the weeks before the Super Bowl and Cinco de Mayo, exactly when Mexican supply tightens. The Dominican Republic plays a different game entirely. About 80% of its U.S. shipments are non-Hass varieties, the larger green-skin fruit that competes for a different shelf rather than going head-to-head with Hass. And Chile, once the giant, now ships about 22 million pounds a year to the U.S., mostly into West Coast retail at the margins of the season. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It pulled the data, built the chart in Python, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: U.S. Census Bureau, USA Trade in Goods, harmonized tariff codes 0804.40.0040 (Hass) and 0804.40.0090 (Other), 2025 calendar year. The full dataset used for this chart is available here. --- ## No horse has broken Secretariat's 1973 Kentucky Derby record URL: https://www.randalolson.com/2026/05/04/kentucky-derby-winning-times-secretariat/ Published: 2026-05-04 Categories: data visualization Tags: beautiful-charts-with-ai, kentucky derby, horse racing, sports data, secretariat Winning times at the Kentucky Derby improved steadily from 1896 through 1973, then hit a wall. Secretariat's 1:59.40 remains the only sub-2:00 Derby ever run. Part of Teaching an AI Agent to Make Beautiful Charts Golden Tempo crossed the finish line Saturday at Churchill Downs in 2:02.27, becoming the 2026 Kentucky Derby champion under first-time female trainer Cherie DeVaux. But look at the time. 2 minutes, 2 seconds. The record set by Secretariat in 1973, 1:59.40, was not touched. It hasn't been for 53 years. From 2:08 to 2:00: 70 years of steady improvement The Kentucky Derby ran at 1.5 miles from its founding in 1875 through 1895. In 1896 the distance was shortened to 1.25 miles, and the modern time series begins there. The first winner at the new distance, Ben Brush, ran 2:07.75. What followed was a long decline in times. Better breeding programs, improved track surfaces, and more refined training methods all pulled times lower across the first half of the 20th century. By the 1930s, winners were regularly in the 2:03 to 2:05 range. By the early 1960s, they were approaching 2:00. In 1964, Northern Dancer ran 2:00.00 flat, arriving at the barrier but not crossing it. Over roughly 70 years, the winning time dropped by about 8 seconds. What Secretariat did in 1973 Secretariat's 1:59.40 on May 5, 1973 was the first sub-2:00 Derby ever run. It remains the only one, with one near-miss: Monarchos in 2001 finished in 1:59.97, the closest any horse has come in the 53 years since. What separated Secretariat wasn't just the clock. He ran each quarter mile faster than the one before, a pattern almost no horse at any distance sustains through the final stretch. When Dr. Thomas Swerczek performed the necropsy in 1989, he found a heart estimated at roughly 22 pounds, about 2.5 times the average thoroughbred heart of around 8 to 9 pounds. The chambers and valves were structurally normal; the organ was simply larger than any Swerczek had seen. Track conditions mattered too: both Secretariat in 1973 and Monarchos in 2001 ran on fast, dry tracks with no measurable precipitation in the prior 24 hours. 53 years without improvement After 1973, the trend line goes flat. The 11-year moving average settles between 2:01 and 2:03 and stays there across 5 decades. Since 1973, only three horses have broken 2:01: Monarchos (1:59.97 in 2001), Spend A Buck (2:00.20 in 1985), and Authentic (2:00.61 in 2020). The plateau isn't a Derby quirk. Research on thoroughbred performance across elite races finds that winning times showed minimal improvement after the mid-20th century, while human sprint and distance records kept falling for decades more. Part of the explanation is structural: thoroughbreds have generation intervals of around 10 to 11 years and relatively low heritability for racing speed, which limits how fast selective breeding can push race times lower. Horses are also bred to win races, not set records. Trainers and jockeys optimize for relative position in the field, not absolute time on the clock. Golden Tempo and what 2026 confirms Golden Tempo's win was remarkable in its own right. A 23-1 longshot sitting in last place through the far turn, he swept the entire field in the stretch. His trainer, Cherie DeVaux, became the first woman to train a Kentucky Derby winner in the race's 152nd running. His time, 2:02.27, was completely ordinary for the modern era. That's not a knock on the horse. It's just where the Derby has lived for over 50 years. The best thoroughbreds in the world, trained with every modern tool available, consistently run the 1.25-mile Churchill Downs course in about 2:01 to 2:04. The sub-2:00 zone belongs to Secretariat and one exceptional May afternoon in 2001. Nothing since has come close. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It researched the data, built the chart in Python, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: Wikipedia's Kentucky Derby winners table. Times cover the 1.25-mile distance only (1896 to 2026); the 1875 to 1895 races ran 1.5 miles and are excluded for comparability. The full dataset used for this chart is available here. --- ## U.S. women 40+ now have more babies per capita than teens URL: https://www.randalolson.com/2026/05/01/us-births-teens-vs-40plus/ Published: 2026-05-01 Categories: data visualization Tags: beautiful-charts-with-ai, birth rate, fertility, demographics, public health, age, teen pregnancy Teen pregnancy was a defining U.S. public health worry of the 1990s. Having a baby at 40 was rare. 30 years later those positions have flipped. Part of Teaching an AI Agent to Make Beautiful Charts 30 years ago, teen pregnancy was a defining anxiety of U.S. public health. Having a baby at 40 was rare and often medically discouraged. Today those positions have flipped. U.S. women in their 40s now give birth at a higher rate than U.S. teenagers. The rates crossed in 2022, and the gap is widening every year. Teen pregnancy stopped making headlines In the early 1990s, the teen birth rate hit a modern peak that public health agencies treated as a generational crisis. Then it started falling, almost every year, for 30 years. By 2025, it had dropped by more than 80%, one of the steepest sustained declines in U.S. public health data. The reasons are unglamorous but well documented. Long-acting reversible contraceptives like IUDs and implants went from rare to routine for the teens who wanted them. Their use rose more than 15-fold between 2005 and 2013, according to the CDC. Sex education improved. Teens started having sex later. The CDC's National Survey of Family Growth shows the share of female teens who had ever had sex fell from 51% in 1988 to 42% by 2017. None of this happened in a single dramatic moment. It accumulated, year after year, until the social problem that had once defined U.S. public health policy had largely solved itself. Older motherhood became ordinary The other half of the chart tells a different story. Women started families later, and they kept doing so. The mean age at first birth in the U.S. rose from 24.9 in 2000 to 27.5 in 2023, pushing more first births past 35 and a meaningful share past 40. At the same time, reproductive medicine moved from boutique to mainstream care. IVF clinics opened in every major metro. Egg freezing crossed a key threshold in 2012, when the American Society for Reproductive Medicine officially dropped its "experimental" label, and uptake climbed sharply afterward. The CDC counted more than 97,000 ART-conceived infants in 2021, more than double the 2003 count. Becoming a mom at 40, 42, or 45 is no longer the exception. It is supported by an entire medical industry that scarcely existed a generation ago. The lines crossed in 2022, and they haven't come back For the entire span of modern U.S. vital statistics, going back to 1933, teen births had outnumbered births to women 40 and older, often by a wide margin. The 2022 crossover was a first. What's happened since is more telling than the moment itself. The gap has widened every year. The chart shows two arcs that came from opposite directions, met in 2022, and haven't reversed. This crossover sits inside a larger fertility shift. In the U.S., the general fertility rate hit a record low in 2025, and the total fertility rate has been below replacement since 2009. I wrote about that macro picture and what it means for U.S. population growth in another post. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It researched the data, built the chart in Python, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data sources: CDC National Vital Statistics Reports, Vol. 74, No. 3 (Driscoll & Hamilton, March 2025), Table 2, for 1990 to 2023; CDC Vital Statistics Rapid Release Report No. 43 (Hamilton, Osterman & Gregory, April 2026), Table 1, for 2024 final and 2025 provisional rates. The 40+ series sums the CDC sub-rates for women 40-44 and women 45 and older (each per 1,000 women in age group). The 2025 figures are provisional and based on 99.95% of birth records received and processed by the National Center for Health Statistics as of February 3, 2026. The full dataset used for this chart is available here. --- ## U.S. drug overdose deaths are down 35% from the 2022 peak URL: https://www.randalolson.com/2026/04/29/us-drug-overdose-deaths-1999-2025/ Published: 2026-04-29 Categories: data visualization Tags: beautiful-charts-with-ai, opioid crisis, drug overdose, public health, fentanyl U.S. drug overdose deaths peaked at 107,941 in 2022, then fell two years in a row to a provisional 70,000 in 2025: a 35% drop from the peak. Part of Teaching an AI Agent to Make Beautiful Charts U.S. drug overdose deaths went up nearly every year from 1999 to 2022, climbing from 16,849 to a peak of 107,941. At the peak, that was more Americans dying of overdoses in a single year than from car crashes and gun violence combined. Then the curve broke. CDC's final 2024 count came in at 79,384, a 26% drop in the age-adjusted death rate and the largest single-year decline CDC has tracked over the past decade. Provisional 2025 data puts the count near 70,000, about 35% below the 2022 peak. That's the first sustained two-year drop in 25 years of data. The chart below shows the full arc, including the three opioid waves that drove the rise. Wave 1: OxyContin flooded the market The FDA approved OxyContin in December 1995, and it went to market in 1996. Purdue Pharma's original label included a claim that the drug's "delayed absorption...is believed to reduce the abuse liability," a statement with no clinical evidence behind it. The FDA removed the claim in 2001 after a GAO inquiry, and added a black-box warning the same year. By then, OxyContin prescriptions had jumped from 316,000 in 1996 to more than 14 million for 2001 and 2002 combined. A 2020 review in the AMA Journal of Ethics argued the deeper failure was the FDA's decision to approve OxyContin with a broad indication that let Purdue market it for everyday chronic pain. Opioid-involved deaths tripled over those same years, from 8,050 in 1999 to 21,089 in 2010. The drugs were reaching people far beyond the severe-pain patients the label was written for, and the pipeline kept feeding the supply. Prescription opioid deaths climbed through 2011 before leveling off, but the people who had become dependent on those drugs did not disappear when prescribing tightened. Wave 2: Heroin filled the gap Around 2010, Purdue reformulated OxyContin to resist crushing and dissolving, cutting off a common route of misuse. Prescribers also began writing fewer opioid scripts under pressure from state and federal regulators. People who had developed opioid dependence needed a substitute, and traffickers were ready. Heroin was available and, by that point, considerably cheaper than diverted pills on the black market. Heroin deaths jumped from 3,036 in 2010 to a peak of 15,482 in 2017. The number of past-year heroin users in the U.S. rose from 404,000 in 2002 to 681,000 in 2013 per SAMHSA's National Survey on Drug Use and Health, a shift that tracked closely with the prescription opioid crackdown. The new heroin users were mostly people already dependent on prescription opioids, switching when pills got harder to get. Wave 3: Fentanyl reshaped the supply Illicitly manufactured fentanyl began entering the U.S. drug supply around 2013. It is 50 to 100 times more potent than morphine, which meant traffickers could ship enormous quantities of opioid activity in a package small enough to fit in an envelope. Drug markets adapted fast. Law enforcement fentanyl seizures increased 426% from 2013 to 2014 alone, and synthetic opioid deaths climbed from 3,105 in 2013 to 73,838 in 2022. By 2016, fentanyl had surpassed heroin as the leading driver of overdose deaths in the U.S. The reason was economic. Fentanyl is inexpensive to produce and extraordinarily potent, so it displaced heroin the way cheap imports displace domestic goods. Heroin deaths dropped steadily after 2017 because the product had changed, not because fewer people were using opioids. Per-use risk stayed high. Fentanyl's margin between a dose that produces a high and one that stops breathing is far tighter than heroin's. COVID drove the steepest single-year jump Overdose deaths had been rising before the pandemic, but 2020 was different. Deaths jumped 30% in a single year, from 70,630 to 91,799. Treatment access collapsed. Clinics closed or cut capacity, court-mandated programs paused, and the in-person support networks people in recovery rely on disappeared overnight. The Commonwealth Fund pointed to economic disruption, social isolation, and reduced access to treatment as the main drivers. The 2020 surge accelerated a trajectory that was already steep. Deaths continued rising through 2021 and 2022, passing 100,000 for the first time in 2021. What drove the two-year decline Researchers disagree on the dominant cause. A January 2026 Science paper by Vangelov, Humphreys, Caulkins, and colleagues argued that disruptions to the fentanyl precursor supply chain out of China after a November 2023 U.S.-China summit cut potency and availability across North American drug markets, and that this supply shock did most of the work. Other researchers, including Nabarun Dasgupta at UNC, credit community-level factors, especially the explosion of community naloxone distribution. Several harm reduction programs scaled substantially over the same period. Naloxone access expanded sharply after the FDA approved an over-the-counter version in 2023, every state now has some form of Good Samaritan overdose law, and access to buprenorphine and methadone broadened through pandemic-era telehealth changes. The honest summary is that supply-side and harm reduction explanations both fit the data, and there is no consensus on the split. The decline carried into 2025, but the pace slowed. Provisional CDC data through November 2025 put the 12-month death count near 70,000, about 12% below the 2024 final total. Humphreys warned in mid-2025 that the slowdown could mean the 2024 drop was a one-off rather than a fundamental change in the epidemic. What the headline number leaves out A 35% drop is real, but several pieces of the picture do not move with it. The "deaths are down" story is overwhelmingly an opioid story. Methamphetamine-involved overdose deaths climbed from 2,266 in 2011 to 34,855 in 2023, and cocaine-involved deaths from 4,681 to 29,449 over the same period, according to CDC MMWR. Both fell in 2024 along with opioid deaths, but stimulant-involved deaths remain near record highs and the share of overdoses involving multiple drugs keeps rising. Racial disparities are widening. Overdose death rates among Black and American Indian / Alaska Native populations have continued to rise or fall more slowly than rates among white Americans, Stateline reported. The gains have not been distributed evenly. A new sedative is moving through the fentanyl supply. Medetomidine, a veterinary tranquilizer whose sedation cannot be reversed by naloxone, has spread fast. CDC reports of medetomidine in seized drug samples rose from 247 in 2023 to 8,233 in 2025, and across CDC sentinel testing sites in late 2025 it was detected in roughly 35% of opioid-positive samples on average, with regional pockets above 50%. The CDC issued a Health Alert Network advisory on it in April 2026. Federal funding for the programs being credited is unstable. On January 14, 2026, the Trump administration sent termination letters wiping out roughly $1.9 billion in SAMHSA addiction and mental health grants, then reversed the cuts within 24 hours after public outcry. Broader cuts to HHS grant funding, Medicaid, and CDC data systems are still on the table. The Brookings Institution flagged the trajectory as a real threat to the durability of the trend. Even after the 35% drop, roughly 70,000 Americans still died of an overdose in 2025, more than double the count from a decade earlier. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It researched the data, built the chart in Python, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: CDC NCHS Data Briefs 329, 522, and 549, covering final overdose death counts 1999 through 2024. The 2025 provisional data point uses the CDC VSRR predicted 12-month count through November 2025, available at data.cdc.gov. The full dataset used for this chart is available here. --- ## U.S. birth rate fell to a record low in 2025 URL: https://www.randalolson.com/2026/04/28/us-birth-rate-record-low-2025/ Published: 2026-04-28 Updated: 2026-06-20 Categories: data visualization Tags: beautiful-charts-with-ai, birth rate, fertility, demographics, public health The CDC reported in April 2026 that the U.S. general fertility rate fell to 53.1 in 2025, a record low in data going back to 1909. Part of Teaching an AI Agent to Make Beautiful Charts On April 9, 2026, the CDC released provisional birth statistics for 2025. The general fertility rate, which measures births per 1,000 women ages 15 to 44, came in at 53.1. That is lower than any year in the dataset going back to 1909, and 23% below the 2007 peak of 69.3. The chart below puts 116 years of that decline in context. Birth rates were already falling when the Depression hit The 1909 rate of 126.8 was not a peak. Births per woman had been declining through the late 1800s as the U.S. shifted from a rural farming economy, where children were a labor asset, to an urban industrial one, where they were an expense. By 1929, the rate had already fallen to about 89. The Depression then cratered it further. Research published in the Journal of Demographic Economics estimated the economic hardship of the 1930s produced roughly 750,000 fewer births nationwide, driven by delayed marriages and a steep drop in marital fertility rates. The rate hit 76.3 in 1933 before slowly recovering through the late 1930s. Returning soldiers triggered the baby boom The GFR climbed steadily from 1940 through 1957, peaking at 122.9. The causes were straightforward: 16 million service members came home after World War II, many of them in their prime childbearing years, and the GI Bill gave them access to low-interest home loans, college tuition, and unemployment benefits. Economic security and deferred family plans collided at the same moment. At the 1957 peak, a baby was born in the U.S. roughly every 7 seconds. The baby boom is often remembered as a new baseline for U.S. family size, but the chart makes clear it was a temporary interruption of a longer downward trend, not a permanent reset. The birth control pill drove the post-boom collapse The FDA approved oral contraceptives in 1960. Rates dropped quickly through the 1960s. The 1965 Supreme Court ruling in Griswold v. Connecticut made contraception legal for married couples; Eisenstadt v. Baird extended the same right to unmarried couples in 1972. Women entered the workforce in larger numbers, and average family size shrank. By 1976, the GFR had fallen to 65.0, nearly half the baby boom peak. Rates then stabilized through the 1980s and 1990s, lifted partly by the echo boom as baby boomers' own children reached childbearing age. By 2007, the rate had edged back up to 69.3, briefly bringing the total fertility rate near the 2.1 replacement level for the first time since the early 1970s. The 2007 decline never turned around The Great Recession hit in 2008, and birth rates fell with it. That part is predictable. What is less predictable is that they kept falling after the recession ended. A 2022 analysis in the Journal of Economic Perspectives called it a puzzle: researchers could not identify any single economic, policy, or social shift after 2007 that explained the continued decline. Student debt, housing costs, and childcare prices each correlated loosely with birth rates in some states but not others, and none could explain a decline that cut across income levels and education groups. The leading explanation is a shift in preferences, particularly among younger cohorts who report lower desired family sizes than their predecessors did at the same age. The 2025 rate of 53.1 is 23% below 2007 and lower than any year in the CDC's dataset, according to the provisional CDC report. What the numbers mean going forward A general fertility rate of 53.1 corresponds to a total fertility rate well below the 2.1 replacement threshold, which is roughly where it needs to be for a population to sustain itself without net immigration. The U.S. TFR has been below replacement consistently since 2009. Social Security's trustees now project the TFR will not return to 1.9 until 2050, a decade later than their prior estimate. Net international migration is on track to become the sole driver of U.S. population growth after 2030, according to the Congressional Budget Office's 2026 demographic outlook. What happens to that migration flow over the next decade will shape the workforce, the tax base, and Social Security's solvency in ways that dwarf anything birth rates alone can fix in the near term. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It researched the data, built the chart in Python, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: National Center for Health Statistics, CDC/NCHS Births and General Fertility Rates dataset, supplemented with provisional data from VSRR Report No. 43 (April 2026). The full dataset used for this chart is available here. --- ## Artemis II ended humanity's 54-year confinement to low Earth orbit URL: https://www.randalolson.com/2026/04/27/human-spaceflight-distance-1961-2026/ Published: 2026-04-27 Categories: data visualization Tags: beautiful-charts-with-ai, space, NASA, Apollo, Artemis After Apollo 17 in 1972, every crewed mission stayed inside low Earth orbit. Artemis II broke that streak in April 2026 on a lunar flyby of 406,771 km. Part of Teaching an AI Agent to Make Beautiful Charts On April 10, 2026, 4 astronauts splashed down in the Pacific after flying farther from Earth than any humans on record. The Artemis II crew reached 406,771 km on a free-return loop around the Moon, eclipsing Apollo 13's 400,171 km from April 1970. The chart below places every crewed spaceflight since Yuri Gagarin in context: year on the horizontal axis, maximum distance from Earth on the vertical axis (log scale). The Apollo window: 9 missions in 5 years Between December 1968 and December 1972, 9 crewed missions reached lunar distances. 8 orbited the Moon, and Apollo 13 looped around after an oxygen tank explosion forced the crew to abort the landing. On the chart, the Apollo missions form a tight cluster near 400,000 km, about 1,000 times higher than the ISS orbits today. Apollo ended because the goal was met and the money ran out. Once Neil Armstrong and Buzz Aldrin walked on the Moon in July 1969, the political rationale for racing there dissolved. NASA's budget had been falling since 1966, and the Nixon administration redirected funds toward Skylab and the Space Shuttle. Apollos 18, 19, and 20 were canceled. Apollo 17 launched in December 1972, and no human left Earth orbit again until April 2026. Which Apollo mission holds the distance record? The official answer is Apollo 13. NASA's April 2026 press release credited Apollo 13's 400,171 km as the mark Artemis II broke, and Guinness World Records lists the same mission. Artemis II exceeded that distance at 12:56 p.m. CDT on April 6, 2026, Flight Day 6. The measurement geometry makes this question harder than it looks. The Earth-Moon distance varies by roughly 43,000 km over the course of a year, and which point on Earth you measure from shifts the result by another 12,000 km. The AIAA Houston paper on this question walks through the calculation in detail. Either way, Artemis II flew past all of them. 54 years at the wall After Apollo 17, every crewed spaceflight for 54 years stayed inside low Earth orbit. The ISS orbits at about 400 km. The Space Shuttle's highest flights were the Hubble deployment and servicing missions, which reached around 600 km. The dense stripe of dots between 300 and 600 km represents the bulk of those 403 post-Apollo missions, from Salyut to Mir to decades of ISS rotations, all confined to the same narrow band. Polaris Dawn pushed the LEO ceiling to 1,400 km in 2024 (more on that below), but no crew left orbital space entirely until Artemis II. The dark middle band on the chart, from 2,500 km to 360,000 km, has 0 dots. That gap is not a data error. No crewed mission has achieved orbit in that range. For 54 years, human spaceflight had a hard ceiling at a few hundred kilometers, and the chart makes that compression visible in a way that a timeline alone cannot. 2 altitude records, 58 years apart 2 missions sit visually apart from the LEO cluster without reaching the Moon: Gemini 11 in 1966 and Polaris Dawn in 2024, both visible as isolated dots above 1,000 km in the blue LEO zone. In September 1966, Pete Conrad and Dick Gordon used the Agena Target Vehicle's engine to boost Gemini 11 to 1,374 km, the highest Earth orbit ever flown. The Shuttle had no reason to go that high, the ISS sits at 400 km, and the record simply persisted for 58 years. In September 2024, SpaceX's Polaris Dawn crew flew to 1,400.7 km to study Van Allen radiation effects on the human body, finally breaking the mark. Both missions appear in the chart as isolated dots at nearly the same height, separated by 58 years on the x-axis. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It researched the data, built the chart in Python, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: List of human spaceflights on Wikipedia (412 missions through Artemis II, April 2026). Apollo lunar distances from NASA mission records; Artemis II distance from the NASA press release, April 6, 2026. --- ## Sabastian Sawe ran the first sub-2-hour marathon: 118 years of marathon records URL: https://www.randalolson.com/2026/04/26/mens-marathon-world-record-1908-2026/ Published: 2026-04-26 Updated: 2026-06-20 Categories: data visualization Tags: beautiful-charts-with-ai, marathon, running, sports, world record On April 26, 2026, Kenya's Sabastian Sawe ran 1:59:30 at London to become the first man to break two hours in a record-eligible marathon. Part of Teaching an AI Agent to Make Beautiful Charts This morning in London, Kenya's Sabastian Sawe ran the marathon in 1 hour, 59 minutes, 30 seconds. He took 65 seconds off Kelvin Kiptum's 2023 world record, and became the first man to break two hours in an open, record-eligible marathon. The chart below puts that performance in the context of the men's world record line going back to the first Olympic marathon in 1908. The 1908 Olympics set both the distance and the first record Johnny Hayes of New York won the 1908 London Olympic marathon in 2:55:18, the first official world best at the modern marathon distance. He inherited the gold from Italy's Dorando Pietri, who finished first but was disqualified after officials propped him up and helped him across the line in front of the Royal Box. The 26 mile, 385 yard distance came from a royal request, not a measured course. Princess Mary asked for the race to start on the East Lawn of Windsor Castle so the royal children could watch from the nursery window, and the finish was placed in front of the Royal Box at the new White City stadium. The exact distance between those two points became the marathon, and the IAAF (now World Athletics) formally adopted it as the standard in 1921. 1953: Jim Peters runs the first sub-2:20 The first half-century of the record looks slow on the chart because almost no one was racing the marathon as a serious event. Most early world bests came from a handful of one-off long-distance races and Olympic finals. The pace of progress accelerated after World War II, and the runner who carried it was an unsponsored optician from Essex named Jim Peters. Peters set the world record four times between 1952 and 1954. His 2:18:40 at the Polytechnic Marathon in June 1953 was the first marathon under 2:20, and he kept lowering his own mark, ending with 2:17:39 a year later. The 2:17:39 stood as the world best for four years. Peters never won an Olympic medal. He collapsed within sight of the finish at the 1954 Empire Games and never raced a marathon again. 1960: Abebe Bikila wins barefoot in Rome Abebe Bikila of Ethiopia ran the 1960 Rome Olympics in 2:15:16 without shoes. The popular story is that this was a gesture; the actual reason was simpler. Adidas, the official Olympic shoe supplier, had run out of shoes that fit him by the time he arrived to collect his pair. The shoes he was given gave him blisters in training, so on race day he left them off. Bikila was the first sub-Saharan African to win an Olympic gold medal. He defended his title four years later in Tokyo, this time in shoes, and broke the world record again with a 2:12:11. The chart's annotation points to the 1960 run because that was the moment African distance running stepped onto the world stage; the dominance has not stopped since. 1988: Dinsamo's record holds for a decade Belayneh Dinsamo of Ethiopia ran 2:06:50 at Rotterdam in April 1988. The record stood for 10 years, 5 months, and 3 days before Ronaldo da Costa of Brazil finally took it at Berlin in 1998. That decade is the longest gap between world records since the 1930s. The 1990s saw a generation of runners chase the mark and miss; East African road racing depth was exploding, but no one could put a single race together at the right pace on the right course. Da Costa's 2:06:05 in Berlin opened the modern era of the marathon record. Of the next 10 men to break it, eight were from Kenya or Ethiopia (Khalid Khannouchi of Morocco and the U.S. is the only other name on the list). The race that broke it most often was Berlin, a flat, fast, late-season course that has seen nine world records since 1998. The Vienna asterisk: Kipchoge 1:59:40 in 2019 The 2:00:00 line on the chart was nearly crossed seven years ago. On October 12, 2019, in a Vienna park, Eliud Kipchoge of Kenya ran 1:59:40 with a rotating phalanx of 41 pacemakers, a pace car projecting a green laser line on the road, and hydration handed off from a moving bicycle. The event, called the INEOS 1:59 Challenge, was designed entirely around producing a sub-2 time. World Athletics did not ratify the run as a record because it was not an open competition. Guinness recognized it as a category of its own. The performance proved the human body could cover 26.2 miles in under two hours; what it did not prove was that anyone could do it in a normal race against other runners, with normal hydration and no escort. That is what stayed open for the next seven years. The super-shoe era The line on the chart drops sharply after 2017 for a reason. Nike released the Vaporfly 4% that year, the first carbon-plated, foam-stacked racing shoe. Researchers at the University of Colorado measured a 4% improvement in running economy in 18 sub-elite runners; the study was Nike-funded but the methodology held up under independent replication. Kipchoge ran 2:01:39 in the Vaporfly at Berlin in 2018. Kiptum ran 2:00:35 in a Nike Alphafly at Chicago in 2023. Every world record in the men's marathon since 2017 has been set in a carbon-plated shoe. Every major shoe brand now sells one. The shoes are not the only reason the record has fallen this fast, but the chart's slope from 2017 forward is steeper than anything in the previous 40 years. Kelvin Kiptum should have been the one The runner most expected to break two hours was Kelvin Kiptum. He debuted at Valencia in December 2022 in 2:01:53, the fastest debut marathon in history. He won London in April 2023 in 2:01:25, then set the world record at Chicago that October in 2:00:35, at age 23. He was scheduled to attempt sub-2 at Rotterdam in April 2024. He never ran it. Kiptum died on February 11, 2024, in a car crash near Kaptagat, Kenya, alongside his coach Gervais Hakizimana. World Athletics had ratified his world record five days earlier. His record stood for two and a half years, longer than the careers of most elite marathon runners last. April 26, 2026: the barrier finally falls Sawe came through halfway in 1:00:29 with Yomif Kejelcha of Ethiopia, both running under world record pace. Sawe broke away at 30K, ran his next 5K in 13:54, and covered the final 2.195K in 5:51, ten seconds faster than anyone has ever closed a marathon. His average pace was 4:33 per mile, for 26.2 miles in a row. The course in London is flat, the conditions were dry and cool, and the field was deep enough that Kejelcha finished second in 1:59:41 to become the second man under two hours. What changes from here is the framing. The 2:00:00 line was the last round-number barrier in the marathon, the running equivalent of Roger Bannister breaking 4:00 in the mile. Bannister ran his 3:59.4 on a cinder track at Oxford in May 1954; 46 days later, John Landy ran 3:58.0. Within three and a half years, Derek Ibbotson had taken the mile world record down to 3:57.2. The pattern after a barrier falls is usually that the next runner gets there faster. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It researched the data, built the chart in Python, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: Marathon world record progression on Wikipedia, drawing on World Athletics and the Association of Road Racing Statisticians. The full dataset is available here. --- ## Every Atlantic hurricane track, 1980 to 2025 URL: https://www.randalolson.com/2026/04/23/atlantic-hurricane-tracks-1980-2025/ Published: 2026-04-23 Categories: data visualization Tags: beautiful-charts-with-ai, weather, hurricanes, climate Every Atlantic tropical cyclone from 1980 to 2025 on one map. The shape of the basin and the deadliest storms of 46 seasons emerge from the density. Part of Teaching an AI Agent to Make Beautiful Charts Over the last 46 years the Atlantic has produced 738 named tropical cyclones. Every one of them has a track: a string of latitudes and longitudes logged at 6-hour intervals by the National Hurricane Center from the moment a storm is named until it dissipates. Plot all 738 on one map, color each segment by sustained wind, and the climate geography of the basin falls out of the mesh. Hurricanes go where the warm water is The tracks cluster in a few specific places: the Gulf of Mexico, the arc of the Caribbean, the stretch of open ocean east of the Lesser Antilles, and the sweep up the U.S. East Coast. The pattern is not random. Tropical cyclones form and intensify only over warm, deep ocean water, and the basin's warm water sits in a specific, repeatable configuration. The chart is effectively a 46-year average of that geography, drawn by every storm that used it. The Gulf and the Caribbean are intensification chambers Notice where the yellow, orange, and red pixels concentrate: inside the Gulf of Mexico, and along a stripe running from the northwestern Caribbean up toward the Bahamas. Warm, deep water with no continental shelf to disrupt it is the fuel a mature hurricane wants. Katrina (2005), Harvey (2017), Michael (2018), Ian (2022), and Milton (2024) all hit their peak intensities inside the Gulf. Wilma (2005) did the same in the northwestern Caribbean, and Hurricane Maria went from a Category 1 to a Category 5 in 24 hours in the eastern Caribbean before striking Puerto Rico in 2017. Two storms share the modern record The chart shows 26 Category 5 storms in 46 years. Two are tied for the strongest sustained winds measured in that span: Allen in 1980 and Melissa in 2025, both at 165 knots (about 190 mph). Allen ran the full Cape Verde route, hit that 165-knot mark three separate times between the Caribbean and the Yucatán Channel, and came ashore in south Texas. Melissa intensified just south of Jamaica and made landfall on the island as a Category 5, the strongest hurricane ever to strike Jamaica. On the chart, both tracks read as the thin red-into-white filaments that stand out against the blue majority. The basin has 4 gears Four patterns are easy to pick out in the mesh. Cape Verde storms form off the coast of West Africa and trace the long westward route across the Atlantic, usually curving north as the mid-latitude westerlies catch them. Gulf storms form closer to home, tend to be shorter-lived, and are often the ones that peak just before landfall. Caribbean storms have the shortest fuse: from formation to landfall can be 48 hours. The fourth gear is recurvature, the sweep of tracks that cross the southern U.S. coast or miss it, get caught in the westerlies, and head northeast along the East Coast before vanishing into the North Atlantic. Those are the storms that threaten the mid-Atlantic and New England. How this chart was made An AI agent built the chart end to end: parsing the HURDAT2 text format, filtering to the satellite era (1980 onward), rendering 738 storm tracks colored by sustained wind speed, and drafting the post. The work was iterated against the Tufte test, a data visualization quality standard from Goodeye Labs. For this piece the goal was beauty first, so the form privileges density and depth over strict minimalism. The source is the National Hurricane Center's HURDAT2 Atlantic basin best-track database, the authoritative record of Atlantic tropical cyclones maintained by the NHC and reprocessed by Chris Landsea's team. The cleaned CSV (21,335 track points across 738 storms) is available: atlantic-hurricane-tracks-1980-2025.csv. The workflow behind it is public: run the same high-signal chart workflow to make your own. --- ## 130 years of Boston Marathon winning times URL: https://www.randalolson.com/2026/04/23/boston-marathon-winning-times-1897-2026/ Published: 2026-04-23 Categories: data visualization Tags: beautiful-charts-with-ai, boston marathon, running, sports John Korir broke a 15-year-old Boston Marathon course record in 2026. The full 1897-2026 arc shows why, and what the women's line has been doing. Part of Teaching an AI Agent to Make Beautiful Charts On Monday, Kenya's John Korir ran 2:01:52 at the Boston Marathon and broke a course record that had stood for 15 years. The previous mark, Geoffrey Mutai's 2:03:02 in 2011, lasted longer than any other in the modern era. Korir beat it by 70 seconds, and two other runners also finished under the old record. The chart below puts that performance in the context of every winning time since 1897. The first race was 24.5 miles and had 15 runners John McDermott of New York won the first B.A.A. Marathon on April 19, 1897 in 2:55:10. He beat a starting field of 15 runners, of whom only 10 finished. The course ran 24.5 miles from Ashland to the Irvington Oval near Copley Square, shorter than the 26.2 miles that became the marathon standard after the 1908 London Olympics. The race itself was a direct import from the 1896 Athens Olympics. B.A.A. track coach John Graham had just managed the U.S. team in Athens, came home inspired, and organized a Boston version on Patriots' Day. It is now the oldest annual marathon in the world. Women were not allowed to enter until 1972 In 1966, Bobbi Gibb hid in the forsythia bushes near the Hopkinton start, jumped into the pack after the gun, and ran the entire course in 3:21:40. She had written to the B.A.A. for an entry form two months earlier, and race director Will Cloney had rejected her with the claim that women were not physiologically capable of the distance. Gibb beat two-thirds of the men's field that day. Women's entries were not officially sanctioned at Boston until 1972, when Nina Kuscsik won the first official women's race. Once the door opened, the progression was extraordinary. The women's winning time dropped by nearly an hour over the next two decades, from Gibb's 3:21:40 in 1966 to Rosa Mota's 2:24:30 in 1988. The line falls steeper on the chart than the men's line ever has. East Africans transformed the men's race after 1988 Ibrahim Hussein won Boston in 1988 in 2:08:43, edging Tanzania's Juma Ikangaa by one second. He became the first African man to win the race. Since 1991, a Kenyan or Ethiopian has won the men's Boston Marathon every year except three: Lee Bong-ju of South Korea in 2001, Meb Keflezighi of the U.S. in 2014, and Yuki Kawauchi of Japan in 2018. The pattern is not an accident of any single marathon. Peer-reviewed research attributes East African distance-running dominance to a combination of chronic high-altitude training in Kenya's Rift Valley (Iten sits near 8,000 feet), biomechanical efficiency, a training culture built around large groups of world-class athletes, and powerful economic motivation. Boston's prize money alone runs to $150,000 for the winner, plus a $50,000 course-record bonus. Korir walked away with $200,000. The women's course record is faster than most men's winners in history Sharon Lokedi of Kenya won Boston in 2025 in 2:17:22, cutting 2 minutes 37 seconds off the 11-year-old women's course record. She repeated in 2026 with a 2:18:51. Her 2025 record is faster than every men's winning time at Boston from 1897 through 1955. It would still have won the men's race in 70 different years of the race's history, including every year of the first 59 races. The 2026 time is even sharper. 2:18:51 is the exact time Keizo Yamada ran to win the men's race in 1953. The women's line has caught up to where the men's line was in the mid-1950s. Women were still 19 years away from being officially allowed to enter the race at that point. Korir's record and the super shoe era Korir's 2:01:52 is the fifth-fastest marathon ever run. It came 15 years after Mutai's 2:03:02, which itself was run with a strong tailwind and stood as the unofficial world best until regulators changed the rules. The gap between them is almost entirely a super shoe story. Nike introduced the Vaporfly 4% in 2017, and peer-reviewed studies of elite race data have since estimated the shoes are worth between one and three percent. Every major shoe brand now builds its own carbon-plated, foam-stacked race shoe. The top of the field at Boston this year ran in some variant of this technology. Three men broke Mutai's old course record in a single race. Alphonce Simbu of Tanzania finished second in 2:02:47, and Benson Kipruto of Kenya finished third in 2:02:50. Korir's older brother Wesley won Boston in 2012, making them the first pair of brothers to win the race. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It researched the data, built the chart in Python, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: Boston Athletic Association race champions, with cross-verification from the Wikipedia list of winners. The full dataset is available here. --- ## A thousand springs in Kyoto, in one chart URL: https://www.randalolson.com/2026/04/23/kyoto-sakura-1200-years/ Published: 2026-04-23 Categories: data visualization Tags: beautiful-charts-with-ai, climate, japan, cherry blossoms Kyoto's 1,200-year cherry blossom record shows modern decades hold the earliest cluster of blooms ever seen. A warming signal written in flowers. Part of Teaching an AI Agent to Make Beautiful Charts Kyoto has been watching its cherry trees bloom for more than a thousand years. Aristocrats, monks, and emperors wrote the dates down in diaries and court records, and a phenologist named Yasuyuki Aono spent years pulling those mentions together into a single dataset. The result is one of the longest continuous biological records in the world: 838 years of recorded full-flower dates, stretching from 812 CE all the way through the 2026 season. For most of Kyoto's history, the cherry trees bloomed around April 14 Each ray on the chart is one year of observation, placed around the circle chronologically. The luminous band running through the middle is the long-term average: day 104 of the year, or roughly April 14. From the 800s through the 1800s, bloom dates drifted around that ring. Warmer and cooler decades produced clusters that reached in or out, but the center of gravity held steady for over a thousand years. The 1,200-year average and the pre-1900 average differ by less than a day. The last 60 years have no parallel in the record Look at the top of the ring, the arc covering the 20th and 21st centuries. It is magenta in a way no other stretch of the chart is. The average full-flower date since 1950 is April 7, 7 days earlier than the long-term average. The 2020s on their own average March 30, almost 15 days earlier. The shift is not a single warm year or a brief episode; it is a sustained pull on the entire distribution. Individual extreme years have always happened. 1409 produced a March 27 bloom that was the earliest in the medieval record. 1323 delivered a May 4 bloom, the latest ever recorded. What is new is the recent concentration of records: 2023 set an all-time earliest at March 25 (day 84), and 2021 came in one day behind at March 26. Singular extremes used to punctuate a stable baseline. Now the baseline has moved, and the records keep falling. Why this record exists at all The data was preserved because of culture, not science. Cherry blossom viewing, hanami, has been a civic and imperial occasion in Kyoto since at least the 9th century. The date of full flower was logged the way a modern city might track its first snowfall or the opening day of a farmers' market. Aono and Saito combined those scattered court references into a usable phenological series; their 2010 paper reconstructs March temperatures in Kyoto from the bloom dates and finds that modern warming separates cleanly from the pre-industrial baseline. How this chart was made An AI agent built the chart end to end: picking the radial form to match the subject, assembling the data, iterating against the Tufte test, a data visualization quality standard from Goodeye Labs, and drafting the post. For this piece the brief was beauty first, so the form leans more ornamental than strictly Tuftean. The source is the Aono, Kazui, and Saito 1,200-year Kyoto cherry tree flowering record from Osaka Prefecture University, extended through 2026 by Our World in Data. The cleaned year-and-day-of-year CSV is available: kyoto-sakura-1200-years.csv. The workflow behind it is public: run the same high-signal chart workflow to make your own. --- ## U.S. oil: from peak importer to net exporter in one decade URL: https://www.randalolson.com/2026/04/23/us-oil-importer-to-net-exporter-1973-2025/ Published: 2026-04-23 Updated: 2026-06-20 Categories: data visualization Tags: beautiful-charts-with-ai, oil, energy, usa, shale The U.S. imported a record 13.7 MMb/d of oil in 2005. By 2025 it exports 2.8 MMb/d more than it imports. Here is what actually changed. Part of Teaching an AI Agent to Make Beautiful Charts Back in 2014 I wrote a post titled Where the U.S. gets its oil from, pointing out that public opinion wildly overestimated how much oil came from the Middle East. The real top supplier was Canada, and 60% of what the country burned was produced at home. That was the story in 2014. It is not the story now. The U.S. has been a net exporter of petroleum since 2020, and in 2025 it exported 2.8 million barrels per day more than it imported. Peak dependence was 2005, not 1979 The mental picture most people carry of U.S. oil dependence belongs to the 1970s: gas lines, the Arab oil embargo, Carter's cardigan. The actual peak came much later. U.S. petroleum imports hit 13.7 MMb/d in 2005, the highest in recorded history. Net imports peaked the same year at 12.5 MMb/d. The country was buying more foreign oil than at any point before or since, and domestic production had been grinding down for 35 years straight. That was the bottom of a long slope. Domestic crude production fell from 9.6 MMb/d in 1970 to about 5.0 MMb/d in 2008. Everyone assumed that curve kept going down forever. It did not. The shale revolution is what actually changed The reversal in the chart is driven by one thing: domestic production. Combining horizontal drilling with hydraulic fracturing made it economic to extract oil from tight shale formations that older vertical wells could not touch. Horizontal wells that once extended 2,000 feet now routinely exceed 15,000 feet, with dozens of separate fracking stages along each lateral. The technique first scaled in the Barnett Shale in Texas around 2005 for natural gas, then was adapted to oil in the Bakken (North Dakota), Eagle Ford (South Texas), and the Permian Basin (West Texas and New Mexico). Those three plays now produce roughly 90% of U.S. tight oil. Between 2010 and 2019, U.S. shale oil production added more than 7 MMb/d, and by 2025 U.S. crude production hit an all-time record of 13.6 MMb/d, 2.7x the 2008 low. The 2015 law that opened the gates Production alone would not have flipped the trade balance this hard. The exports line on the chart stays near zero for four decades because it had to. The Energy Policy and Conservation Act of 1975 banned most crude oil exports in response to the Arab oil embargo, and the ban held for 40 years. Refined products (gasoline, diesel, jet fuel) could leave, but raw crude could not. That changed on December 18, 2015, when President Obama signed the Consolidated Appropriations Act, which repealed the ban as Section 101 of Division O. A week later, the Commerce Department confirmed that "a license is no longer required to export crude oil" from the United States. U.S. crude oil exports went from 0.5 MMb/d in 2015 to 4.1 MMb/d in 2024. That surge is the second half of the chart. 2020: the crossover The two lines cross for the first time in 2020. COVID demand destruction helped: the pandemic flattened import volumes and pulled global crude prices negative for a few hours in April. But the trajectory was already obvious before the lockdowns, and it held after them. Every year since 2020 has been a net-export year, and 2025 set a record at 2.8 MMb/d of net exports. The U.S. still imports a lot of oil: 7.9 MMb/d in 2025, mostly heavy crude that Gulf Coast refineries were built to process. The composition is much simpler than it was in 2014, though. Canada supplied 62% of U.S. crude oil imports in 2024, a record share driven by the May 2024 completion of the Trans Mountain pipeline expansion from Alberta. Mexico is a distant second, shipping 464,000 b/d of crude in 2024, down 37% from 2023 as Mexican field productivity fell. Venezuela, once a leading supplier, was at roughly 150,000 b/d in late 2025 under intermittent sanctions. What my 2014 post got right, and what it missed I was right that the Middle East was never the main supplier. That misconception was wrong then and is even more wrong now. What I did not anticipate, because the data did not yet show it, was how quickly the whole question would stop mattering. "Where does the U.S. get its oil from" was the right question in 2014. In 2025 the more useful question is "where is all this U.S. oil going?" The answer, per the EIA's 2024 export breakdown, is mostly Europe and Asia. The Netherlands (home to Rotterdam's trading hub) received the most U.S. crude in 2024 at 825,000 b/d, followed by South Korea, Canada, the U.K., Singapore, and India. Europe as a whole took 1.93 MMb/d, a record share after the 2022 sanctions cut Russian barrels off from Western markets. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It pulled the latest EIA data, prototyped three candidate chart types, picked the one that told the reversal story most clearly, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: U.S. Energy Information Administration, series MTTIMUS2 (imports) and MTTEXUS2 (exports), crude oil plus petroleum products, annual. The combined 1973-2025 dataset is available here. --- ## Libraries were transforming. Then COVID hit. URL: https://www.randalolson.com/2026/04/22/us-public-library-transformation/ Published: 2026-04-22 Categories: data visualization Tags: beautiful-charts-with-ai, libraries, usa Two decades of IMLS data show U.S. libraries were successfully pivoting from book warehouses to community hubs. Then the pandemic set the transformation back years. Part of Teaching an AI Agent to Make Beautiful Charts Everyone keeps saying public libraries are dying. I pulled 24 years of federal survey data to find out, and the picture is more complicated than the headlines suggest. Libraries were in the middle of a transformation when COVID interrupted it. The recession filled the libraries In-person visits peaked at 1.6 billion in 2009, right in the teeth of the Great Recession. Libraries became lifelines: free internet, resume workshops, foreclosure seminars. CNN reported at the time that the Buffalo/Erie County system saw computer use jump 50% as unemployed workers flooded in to search for jobs. Then visits started falling and never stopped. By 2019, before anyone had heard of COVID, annual visits were already down 21% from that recession peak. Smartphone ownership crossed 50% in 2013 and now sits at 90%. You no longer need to visit a building to get online. The computer lab emptied out Public computer use peaked at 369 million sessions in 2010, then fell 74% to just 97 million by 2023. No metric on this chart declined more steeply. In Connecticut, library computer sessions dropped from nearly 6 million in 2006 to under 1 million by 2022. The decline started well before the pandemic. Every year from 2010 onward, fewer people used a library computer than the year before. Smartphones didn't just compete with library terminals; they made them obsolete for most of the population. But libraries were reinventing themselves Even as foot traffic fell, libraries nearly doubled their programming between 2006 and 2019, growing from 3 million programs to 5.9 million. Program attendance grew 67% over the same period, from 75 million to 126 million. Libraries were trading visitors for participants. This was a deliberate strategy. As fewer people needed libraries for books or internet access, libraries started offering STEM makerspaces, coding classes, 3D printing labs, and job training. By FY 2017, over 118 million people attended library programs annually. The transformation was working. COVID set it all back The pandemic demolished the new model along with the old one. Programs offered fell 40% in 2020, from 5.9 million to 3.6 million, then bottomed out at 2.0 million in 2021. Attendance followed the same path, from 126 million down to 79 million and then to 41 million. More than a decade of programming gains were erased in two years. By 2023, programs had recovered to 4.6 million and attendance to 94 million. That's a strong rebound, but both are still well below where they were in 2019. The transformation trajectory that was building for over a decade is now rebuilding from a lower baseline. Digital is the one bright spot Digital checkouts are the only metric on this chart above pre-pandemic levels, up 57% since 2019. OverDrive alone reported 662 million digital checkouts in 2023, and that number jumped another 17% to 739 million in 2024. When library buildings closed in 2020, digital checkouts jumped 23% in a single year and have kept climbing every year since. Turns out that once people discover they can borrow ebooks and audiobooks from their couch, most of them keep doing it. The buildings are still there The number of U.S. public library outlets has held steady at roughly 17,000 for over two decades. Despite every trend on this chart, communities aren't closing their libraries. The open question is whether the programming rebound continues, or whether the pandemic permanently lowered the ceiling. Federal IMLS funding is being cut, and at $266.7 million (75 cents per American), it wasn't exactly generous to begin with. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It downloaded 24 years of IMLS survey data, aggregated national totals, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: Institute of Museum and Library Services, Public Libraries Survey, FY 2000-2023. The aggregated dataset is available here. --- ## The Claude Code leak in four charts: half a million lines, three accidents, forty tools URL: https://www.randalolson.com/2026/04/02/claude-code-leak-four-charts/ Published: 2026-04-02 Categories: data visualization Tags: beautiful-charts-with-ai, Claude Code, Anthropic, npm, cybersecurity Community mirrors of the @anthropic-ai/claude-code npm package spilled 513k lines of TypeScript. Four charts break down what was inside. Part of Teaching an AI Agent to Make Beautiful Charts Version 2.1.88 of @anthropic-ai/claude-code hit npm on March 31, 2026 with a production source map still attached. Security researcher Chaofan Shou spotted the exposure; the map pointed at a zip archive on Anthropic’s Cloudflare R2 storage, and decompressing it recovered the full TypeScript tree. Anthropic told reporters the same thing it tells everyone after these slips: packaging mistake, human error, no customer credentials in the bundle. One early GitHub mirror collected over 41,500 forks before the company had time to respond. Mirrors are faster than statements. I pulled one mirror of 2.1.88, counted 512,664 lines across 1,884 TypeScript and TSX files, and split the snapshot four ways below. None of this is an official Anthropic drop. It is a frozen tarball people can agree on while the fork count climbs. The money is in the plumbing, not the keynote slide utils/ is roughly 180k lines, a little over a third of the tree. Add components/, services/, tools/, and commands/, and you have the shape of a serious desktop agent: React-style terminal UI (Ink shows up in writeups on the first leak), service glue, and a thick tool layer where the product actually touches the filesystem and shell. hooks/, bridge/, and cli/ pile on another ~44k lines together: session hooks, IDE bridge glue, and the entry wiring that turns “agent in your repo” from a slogan into integrations. The Register pegged the March drop at about 512k lines in ~1,900 files, same ballpark as this count. The flag-gated directories are where the rumor threads live. In this mirror memdir/ edges out buddy/ on raw line count; coordinator/ and voice/ sit in the low thousands inside a half-million-line tree. Great fuel for speculation, useless for sizing what actually ships. Three accidents, no villain, escalating stakes February 2025 was the dress rehearsal. Developer Dave Schumaker opened an early cli.mjs, spotted a fat inline sourceMappingURL, went to a vet appointment, and came back to find Anthropic had already yanked the string in a patch. He recovered the map from an undo buffer in Sublime Text. Eighteen million characters of base64 in one line, npm cache archaeology, undo history as forensics. That is not a hack story. It is a release-engineering mistake that repeated on a bigger stage. Five days before the npm fire drill, the CMS leaked marketing drafts. Fortune reported that Anthropic acknowledged testing a new model after researchers found a misconfigured content store. InfoWorld walked through draft copy describing a phased rollout aimed first at security teams, plus the odd detail that internal naming still said “Capybara” in places. Different failure mode than npm, same theme: internal material treated as public by default. March 31 put the CLI on front pages. The Register reported one early GitHub mirror had already been forked more than 41,500 times before the week was out; an uploader later repurposed his repository into a Python feature port over IP liability concerns while forks kept spreading. InfoWorld quoted security researcher Tanya Janca on why this hurts more than a random npm typo: high-value IP means attackers can skip slow reverse engineering and read logic bugs in plain text. Enterprise SKU, gacha mechanics The buddy/ subtree is a weighted collectible game: eighteen species, five rarity tiers, hats, eye glyphs, stats named like inside jokes. The chart mirrors the probabilities baked into buddy/types.ts (60 / 25 / 10 / 4 / 1). It reads like a side project smuggled into a repo that otherwise worries about JWT bridges and task orchestration. Compile-time gates mean this may never ship as-is, or may ship under another flag. But the rarity weights and stat names are in there now, sitting next to the task orchestrator and the IDE bridge. Some repos have easter eggs. This one has a gacha simulator. Forty tools, one attack surface narrative tools/ is about 50.8k lines in 184 files and forty concrete modules. The chart groups them the way the source does: file IO, shells and REPL, web fetch and search, multi-agent task plumbing, plan and worktree modes, MCP and skills hooks, and a small Special / Internal slice (config, user prompts, sleep, synthetic output). The multi-agent block is worth a second look. Eight tools cover spawning sub-agents, tracking tasks, reading output, and stopping runaway processes. That is not a chatbot with bash access. It is an orchestration platform built to manage parallel agents with independent task contexts. The tool list is also a straightforward capability inventory for anyone trying to understand what “AI agent” means in practice: read, write, execute, search, browse the web, spawn more agents, repeat. If you want to understand what an AI agent can do on your machine, this is the table of contents. How this chart was made An AI agent produced these four matplotlib figures as part of the Beautiful Charts with AI series. Each view was refined against the Tufte Test, a data visualization quality standard from Goodeye Labs. The treemap uses squarify; the timeline uses proportional dates on the top axis and a normalized hour strip below; the buddy and tools panels are layout code driven by tables in the mirrored tree. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: aggregated counts from a community mirror of @anthropic-ai/claude-code@2.1.88 (npm publication March 31, 2026). The directory totals table is available here; the tool module listing is available here. --- ## Every generation rediscovers art: 270 years of cultural interests in one chart URL: https://www.randalolson.com/2026/03/30/rise-and-fall-of-art-forms/ Published: 2026-03-30 Categories: data visualization Tags: beautiful-charts-with-ai, data visualization, art history, cultural trends Engraving dominated the 1800s until photography killed it. 270 years of book data show how every generation picks its own art forms. Part of Teaching an AI Agent to Make Beautiful Charts No art form stays on top forever. When a craft catches fire, people write about it: instruction manuals, exhibition catalogs, critical essays, histories. That published record makes the Google Books Ngram Viewer, which tracks word frequency across millions of English books, a surprisingly good thermometer for cultural interest in art forms over time. Plotting 15 of them across 270 years tells a clear story: each generation finds its own creative obsession, peaks, then fades as something new takes over. Engraving ruled the 1800s, then photography killed it Engraving was the dominant image reproduction method for over a century before photography made it obsolete in a single generation. Before cameras, every illustration in a book, newspaper, or magazine had to be carved by hand into a metal plate or woodblock. Engraving mentions peaked around 1862, the exact period when engravers were still essential but the writing was on the wall. Daguerre had announced his photographic process in 1839, and by 1861, books were already using photography-aided wood engraving. The real death blow came on March 4, 1880, when the New York Daily Graphic published the first halftone photograph in a newspaper. Within 15 years, halftone screens had replaced engraving in most publications. The craft retreated from commercial work into fine-art printmaking, where it remains today as a niche pursuit. The Arts and Crafts Movement tried to turn back the clock Embroidery's peak around 1885 maps almost exactly onto the height of the Arts and Crafts Movement. William Morris founded Morris & Co. in 1861 to produce hand-crafted decorative arts as a deliberate counter to industrial mass production. He considered Victorian machine-made needlework a degradation of the craft, and his firm's embroidery, tapestry, and textile work revived techniques that had been abandoned for decades. The movement had institutional backing, too. The Royal School of Needlework was founded in 1872 specifically to restore "ornamental needlework to the high place it once held amongst decorative arts." Morris's daughter May Morris became the leading practitioner of art embroidery within the movement. But the revival was swimming against the industrial tide, and by the early 1900s, mentions of embroidery had begun their long decline. Two world wars turned knitting into a patriotic duty Knitting spiked around 1919 because governments turned it into a war effort. When the U.S. entered World War I, the American Red Cross set an immediate target: 1.5 million each of knitted wristlets, mufflers, sweaters, and socks. The campaign slogan was "Knit Your Bit." Local chapters distributed standardized patterns and yarn in Army khaki and Navy blue. The scale was hard to believe. Red Cross volunteers produced an estimated 370 million knitted items using 45 million pounds of wool during the 18 months the U.S. was at war. Red Cross membership exploded from 17,000 to over 20 million. The 1919 Ngram peak reflects a publishing lag: the actual knitting frenzy peaked in 1917-18, and books about the campaigns followed a year or two later. Pottery became ceramics, and it wasn't just a name change The data shows two distinct peaks for clay-based art: "pottery" around 1934 and "ceramics" around 1991, and the two peaks tell different stories. The pottery peak coincides with the studio pottery movement launched by Bernard Leach, who co-founded his St Ives pottery with Shoji Hamada in 1920 and published A Potter's Book in 1940, often called "the potter's bible." By the 1970s and 1980s, a new generation of clay artists pushed back against the Leach school's emphasis on functional vessels and Japanese aesthetics. Artists working in sculptural and conceptual modes deliberately adopted "ceramic artist" and "ceramics" over "potter" and "pottery" to claim fine-art status. The Ngram data captured what actually happened in studios and galleries: a generational break with the past. Photography: from zero to cultural dominance in 140 years Photography went from zero mentions before 1839 to the most-discussed art form by the late 1900s. The Kodak Brownie, launched in 1900 at $1, sold 10 million units in five years and put cameras in the hands of people who had never made a picture before. That was when photography stopped being a specialist technology and became a universal cultural practice. Artistic legitimacy came slower. Alfred Stieglitz's 291 gallery in New York (1905-1917) was the first sustained effort to exhibit photographs as fine art alongside paintings. MoMA established its Department of Photography in 1940, the first major art museum to formally recognize the medium. But broad institutional acceptance didn't arrive until the late 1970s, which aligns almost exactly with the Ngram peak around 1978. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. The small multiples format was chosen to show all 15 art forms with individual year axes, so each trajectory is readable without scrolling. Each panel is normalized to its own peak so that smaller art forms (like watercolor) are as visually clear as dominant ones (like photography). The chronological sorting by peak year flows left-to-right, top-to-bottom, revealing the generational wave. The chart was evaluated against the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data comes from the Google Books Ngram Viewer (English corpus, 2019 edition), which tracks word frequency across millions of published books from 1500 to 2019. The raw data is available as a CSV download. --- ## The rise and fall of bowling in the United States URL: https://www.randalolson.com/2026/03/30/the-rise-and-fall-of-bowling-in-the-united-states/ Published: 2026-03-30 Updated: 2026-06-20 Categories: data visualization Tags: beautiful-charts-with-ai, bowling, usa The U.S. had roughly 12,000 bowling alleys at the peak of the bowling boom in the mid-1960s. By 2023, only 3,154 remained. Part of Teaching an AI Agent to Make Beautiful Charts In the mid-1960s, there were roughly 12,000 bowling alleys scattered across the United States. Today, fewer than 3,200 remain. The chart below traces that entire arc, from the post-war bowling boom through six decades of steady decline. The machine that made bowling boom Before the 1950s, bowling alleys relied on "pin boys" to manually reset pins after each throw. The job was dangerous and unreliable. In 1946, AMF unveiled its Automatic Pinspotter at the American Bowling Congress tournament in Buffalo, NY. By 1952 the machines were in commercial production, and they transformed the sport almost overnight. The number of bowling centers nearly doubled in under a decade, jumping from around 6,600 in 1955 to 11,000 by 1963. Automation made bowling quieter and more family-friendly. New alleys popped up in the suburbs with big parking lots and cocktail lounges, earning the nickname "the poor man's country club." Meanwhile, ABC's Pro Bowlers Tour debuted in 1962 and drew 12-14 million weekly viewers at its peak, routinely outdrawing college football. The long decline The boom didn't last. By the late 1970s, league bowling membership had peaked at 9.8 million and began falling. Between 1980 and 1993, league bowling dropped by 40%. Women entering the workforce had less time for weekly league nights. Cable TV exploded the number of sports options competing for attention. Video games, and later the internet, siphoned away the casual recreation dollars that bowling depended on. The real estate math stopped working, too. A bowling alley sits on 20,000 to 40,000 square feet of prime commercial land. As property values climbed, landlords and developers found more profitable uses for those footprints. The industry's own analysts noted that bowling leaders "were misreading the membership and interest trends during the 1980s and did very little to pump new life into a slumping industry." Bowling alone In 2000, Harvard political scientist Robert Putnam made bowling the central metaphor in his book Bowling Alone. His core observation: the total number of Americans who bowled increased by 10% between 1980 and 1998, but league bowling fell by 40%. People were still showing up to bowl. They just weren't joining leagues to do it. Putnam argued that the decline of league bowling was part of a collapse in American civic life. Drawing on nearly 500,000 interviews, he showed that Americans were signing fewer petitions, joining fewer organizations, knowing their neighbors less, and socializing with their own families less often. The share of Americans who believed "most people can be trusted" fell from 55% in 1960 to roughly 35% by the late 1990s. Bowling was the canary in the coal mine. Critics later argued that civic life was shifting forms rather than disappearing, but the bowling numbers themselves were hard to dispute. What 3,154 bowling alleys look like The Census Bureau counted 3,154 bowling establishments with paid employees in 2023, down from 6,148 in 1986, the first year of consistent federal data. That's a 49% decline in under four decades, on top of the roughly 50% decline that had already occurred between the mid-1960s peak and 1986. The surviving industry looks nothing like the one that peaked in the 1960s. Companies like Bowlero (now rebranded as Lucky Strike Entertainment) have pushed upscale: cocktail bars, gourmet food, cosmic bowling with blacklights and thumping music. Food and beverage now accounts for 35-45% of a modern bowling center's revenue, up from a small fraction in the league era. The reinvention has slowed the bleeding, but it hasn't stopped it. The decline rate in 2019-2023 (-2.5% per year) is virtually identical to 2015-2019 (-2.2% per year). How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It researched and compiled the data, built the chart in Python, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data sources: Industry estimates for 1940-1965 compiled from the Smithsonian, Reference for Business, and other historical sources. Federal data for 1986-2023 from the U.S. Census Bureau, County Business Patterns (NAICS 713950 / SIC 7933). The full dataset is available here. --- ## The rise and fall of nuclear weapons testing URL: https://www.randalolson.com/2026/03/27/the-rise-and-fall-of-nuclear-weapons-testing/ Published: 2026-03-27 Categories: data visualization Tags: beautiful-charts-with-ai, nuclear weapons, cold war, geopolitics Between 1945 and 1998, seven countries detonated over 2,000 nuclear weapons across just a handful of remote test sites. This chart maps every one of them. Part of Teaching an AI Agent to Make Beautiful Charts Between 1945 and 1998, seven countries detonated over 2,000 nuclear weapons. Not in wars, but in tests. The chart below plots every single one of them. 1962: the year the world detonated 178 nuclear weapons The peak of nuclear testing came in 1962, with 178 detonations in a single year. To put that in perspective, that's one nuclear explosion every two days, sustained for an entire year. The U.S. conducted 96 tests and the Soviet Union conducted 78. The surge started when the Soviet Union broke a three-year voluntary moratorium in September 1961 with a massive test series that included the Tsar Bomba, a 50-megaton hydrogen bomb and the largest nuclear weapon ever detonated. The U.S. responded with Operation Dominic, a rapid series of 31 atmospheric shots in the Pacific totaling 38 megatons. Both sides were racing to complete as many weapons design validations as possible before the anticipated test ban treaty took effect. Four test sites bore the brunt Over 2,000 nuclear weapons were tested at just a handful of sites. The Nevada Test Site alone hosted 928 tests, roughly 65 miles from downtown Las Vegas. In the 1950s, the mushroom clouds were visible from the Strip. The Las Vegas Chamber of Commerce published calendars listing scheduled detonation times and the best viewing spots. Casinos hosted "dawn bomb parties" for tourists. The Soviet Union split its testing between Semipalatinsk in Kazakhstan (456 tests) and Novaya Zemlya in the Arctic (130 tests). Kazakhstan now recognizes more than one million citizens as victims of Soviet-era radiation exposure from Semipalatinsk. France conducted 193 tests at Mururoa and Fangataufa atolls in French Polynesia, the controversy over which led French intelligence to sink Greenpeace's Rainbow Warrior in Auckland harbor in 1985, killing photographer Fernando Pereira. Strontium-90 in baby teeth ended the atmospheric era The shift from atmospheric to underground testing is one of the clearest patterns in the chart: bold bubbles dominate the left side, then mostly vanish after 1963, replaced by small diamonds. The 1963 Partial Test Ban Treaty virtually eliminated above-ground nuclear detonations, and the reason was radioactive fallout showing up in children. The turning point was the Baby Tooth Survey, a citizen science project launched in 1958 that collected over 320,000 baby teeth from children in St. Louis. Children born in 1963 had strontium-90 levels 50 times higher than children born in 1950. The isotope, a byproduct of atmospheric nuclear blasts, was entering the food chain through contaminated milk. These findings directly influenced President Kennedy's decision to sign the Partial Test Ban Treaty on August 5, 1963. Within five years, strontium-90 levels in the milk supply dropped by more than half. 156 nuclear explosions for "peace" Not all underground tests were weapons. The hollow diamonds scattered across Russia at various latitudes are "peaceful" nuclear explosions. The Soviet Union detonated 156 nuclear devices for industrial purposes between 1965 and 1988: seismic sounding for oil and gas exploration, creating underground storage cavities, extinguishing runaway gas well fires, and even attempting to dig canals. The program was called "Nuclear Explosions for the National Economy," and you can see these tests scattered across dozens of sites far from Semipalatinsk. India borrowed this playbook for its first nuclear test in 1974, codenamed "Smiling Buddha." By calling it a "peaceful nuclear explosion," India technically avoided violating the Non-Proliferation Treaty. Few countries bought the distinction. From 178 tests a year to zero After the 1962 peak, nuclear testing declined steadily for three decades. The Partial Test Ban Treaty pushed all testing underground, but it didn't reduce the total number of tests, and notably France and China refused to sign it. France continued atmospheric testing at Mururoa until 1974. China kept going until 1980, when it conducted the last atmospheric nuclear test by any nation. What actually reduced testing was the end of the Cold War. The Soviet Union conducted its last test in October 1990, the UK followed in November 1991, and the U.S. ran its final test in September 1992. France and China held out until 1996, conducting their last tests just months before signing the Comprehensive Nuclear-Test-Ban Treaty. The data ends at 1998, when India detonated five devices in May and Pakistan responded with six of its own just 15 days later. The only country to test since then is North Korea, which conducted six underground tests between 2006 and 2017. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. Inspired by Minard's 1869 visualization of Napoleon's march, the chart uses a single coordinate space to encode six data dimensions simultaneously: year (x-axis), latitude (y-axis, which naturally groups test sites), explosive yield (bubble size), country (color), test type (atmospheric circles vs. underground diamonds), and purpose (filled diamonds for weapons tests vs. hollow diamonds for "peaceful" nuclear explosions). The AI agent iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: SIPRI / Oklahoma Geological Survey Nuclear Explosion Catalog, compiled by data-is-plural. The full dataset is available here. --- ## The engineering and tech gender gap has barely budged in 50 years URL: https://www.randalolson.com/2026/03/26/bachelors-degrees-women-by-major-1970-2022/ Published: 2026-03-26 Categories: data visualization Tags: beautiful-charts-with-ai, education, gender gap, STEM, women in science, usa Computer Science and Engineering remain below 25% women after 50 years. How the gender composition of U.S. college majors changed from 1970-2022. Part of Teaching an AI Agent to Make Beautiful Charts One of my most-shared charts from 2014 showed the percentage of bachelor's degrees earned by women across every major from 1970 to 2012. That post ended with Computer Science stuck at 18% women after a 30-year slide from its 1984 peak. A decade of new NCES data later, what changed? Computer Science: lost decades, slow recovery CS has an unusual trajectory. Women earned 37% of CS bachelor's degrees in 1984, then the number fell for nearly 30 straight years, bottoming out around 18% in 2010. NPR's Planet Money dug into this in 2014 and found that the decline started right when personal computers showed up in American homes, marketed almost entirely to boys. Researcher Jane Margolis interviewed hundreds of CS students at Carnegie Mellon and found that families were far more likely to put the computer in a son's bedroom than a daughter's. By the time these students got to college, professors assumed everyone had been coding since middle school. The good news: CS has been creeping back up, reaching 23% by 2022. That's the highest share since 1987, but still well below the 37% peak from four decades ago. At this pace of recovery, CS won't return to its 1984 level until the 2040s. Individual schools have shown it's possible to move faster: Harvey Mudd went from 10% to over 50% women CS majors by redesigning its intro CS course and changing how it recruited. But scaling what one small college did to an entire national pipeline is a different problem. Engineering: steady but glacial progress Engineering went from 1% women in 1970 to 24% in 2022. That sounds like progress until you realize the growth has been roughly linear for 50 years. There was no sudden breakthrough, no inflection point where engineering suddenly became welcoming to women. Just a slow, grinding climb of about half a percentage point per year. At 24%, engineering today sits where CS was in 1976. If the linear trend holds, engineering won't reach 37% (CS's 1984 peak) until the late 2040s. The "ET gap" is real, the "STEM gap" is not This was the main takeaway from the 2014 post, and the data only reinforces it. Biology hit 66% women by 2022. Physical Sciences climbed from 14% in 1970 to 45% in 2022. Math held steady around 41-43%. These fields are either at parity or within striking distance. The ones stuck far below are CS and Engineering. The gender gap in STEM is really an Engineering and Technology gap. Title IX passed in 1972, two years into this dataset, and women surpassed men in overall bachelor's degrees by 1982. Most fields responded. CS and Engineering didn't. Business quietly crossed 50% Something often lost in the STEM conversation: Business crossed the 50% line around 2002 but has since slid back to 47% by 2022. It's the largest degree category by total enrollment, so even small percentage shifts represent tens of thousands of students. Business went from 9% women in 1970 to near-parity in one generation. The fields at the top keep climbing Health Professions (85%) and Education (83%) have been majority-female for the entire 52-year span of the data. Psychology took a different path: it was 44% women in 1970, crossed 50% in 1973, and has climbed steadily to 80%, one of the largest shifts in the entire dataset. All three center on direct human service work. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It compiled NCES data across multiple digest years, built the chart in Python, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: NCES Digest of Education Statistics, table 318.30 (academic years 1970-71 through 2021-22). The 2021-22 academic year is the most recent available; NCES publishes with roughly a two-year lag. The compiled dataset is available here. --- ## Housing and getting around eat half of the average American household's spending URL: https://www.randalolson.com/2026/03/26/household-spending-2024/ Published: 2026-03-26 Updated: 2026-06-20 Categories: data visualization Tags: beautiful-charts-with-ai, household spending, consumer expenditure, economics A treemap of BLS data shows where the average American household spent its $78,535 in 2024. Housing and transportation alone account for half. Part of Teaching an AI Agent to Make Beautiful Charts Every year the Bureau of Labor Statistics asks thousands of American households exactly where their money goes. The 2024 Consumer Expenditure Survey came out in December 2025 and puts average annual household spending at $78,535. I mapped every dollar. One in three dollars goes to keeping a roof overhead Housing runs $26,266 per year, a full third of the average household budget. Shelter (rent, mortgage, property tax) alone is $16,317 of that. Nobody who has tried to rent or buy in the last five years needs to be told this. Median gross rent hit $1,487 per month in 2024, up 36% from $1,097 in 2019, based on the ACS 1-year estimates for 2019 and 2024. Median household income rose from $65,712 to $81,604 over the same stretch, a 24% increase. The root cause is a supply shortage. Since 2000, housing construction has fallen 3 to 6 million units behind demand from population growth and an aging population forming more households. Home insurance premiums have jumped 57% since 2019 on top of that. Getting around costs more than eating Transportation is $13,318 per year, 31% more than the $10,169 spent on food. Vehicle purchases ($5,337) and insurance, maintenance, and fees ($4,206) are the big line items. Gas is actually a smaller slice at $2,644. For lower-income households, the burden is even worse: the Bureau of Transportation Statistics found they spend up to 32% of pre-tax income just getting to work. Most American cities were built around cars, and the average household doesn't have much choice but to own and maintain at least one. Dining out is closing in on groceries Americans now spend 39 cents of every food dollar at restaurants or ordering delivery. Eating out cost $3,945 in 2024, compared to $6,224 on groceries. That gap keeps narrowing: since 2010, restaurant and takeout spending has grown at about 7% annually, nearly double the 4% growth rate for groceries. Food delivery spending has grown 924% since 1997, crossing $100 billion for the first time in 2024. DoorDash, Uber Eats, and their competitors have made ordering a meal almost as easy as opening the fridge. If those growth rates hold, dining out will match grocery spending by around 2040. 2.3 times more on entertainment than education Entertainment: $3,609. Education: $1,569. That gap looks stark, but the education number only captures what households pay directly for tuition, fees, and supplies. Employer tuition benefits, state university subsidies, and federal financial aid all reduce the out-of-pocket figure. Entertainment, meanwhile, covers everything from streaming subscriptions and pet care to concert tickets and gym memberships. Healthcare looks smaller here because this chart only counts what households pay directly The $6,197 healthcare total in this chart is household spending, not total healthcare spending on a household's behalf. In the BLS data, that breaks into $4,055 for health insurance and $2,142 for medical services, drugs, and supplies. KFF's 2024 Employer Health Benefits Survey found average employer-sponsored family coverage cost $25,572, with workers contributing $6,296 toward that total. For households with employer coverage, most of the premium is paid outside this chart. National health spending hit nearly $5.3 trillion in 2024, growing 7.2% in a single year. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It pulled the latest BLS data, built a treemap in Python, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: Bureau of Labor Statistics, Consumer Expenditure Survey (2024). The cleaned dataset is available here. --- ## Americans eat 3x more cheese and half as much milk as they did in 1970 URL: https://www.randalolson.com/2026/03/26/how-americas-diet-has-transformed-since-1970/ Published: 2026-03-26 Updated: 2026-06-20 Categories: data visualization Tags: beautiful-charts-with-ai, food, usa USDA data reveals how the American diet has transformed since 1970. Cheese consumption tripled, chicken surpassed beef, and fluid milk was cut in half. Part of Teaching an AI Agent to Make Beautiful Charts The USDA has been tracking what Americans eat since 1909. I dug into their Food Availability Per Capita Data System to see what changed since 1970, and the answer is: almost everything. Americans ditched their milk and tripled their cheese Fluid milk consumption was cut in half, from 269 to 134 pounds per person. Cheese tripled, from 11 to 39 pounds. A USDA study found that the milk decline is about fewer drinking occasions, not smaller glasses. Americans just stopped pouring milk at lunch and dinner. Sodas, bottled water, and eventually plant-based milks took its place. Total dairy consumption actually went up over this period. The dairy industry shifted production toward cheese, where margins are better and the product lasts longer on shelves. The USDA credits the explosion of Italian and Tex-Mex cuisines, frozen pizza, and pre-shredded bags for the demand side. Italian cheese varieties alone went from 2.1 pounds per person in 1970 to nearly 15 pounds by 2012. Not shown on the chart: yogurt rose 1,627%, from under 1 pound per person to 14.3, mostly on the back of the Greek yogurt boom in the late 2000s. Chicken overtook beef in 2010 Chicken more than doubled, from 27 to 68 pounds per person. Beef dropped 30%, from 80 to 56 pounds. According to USDA analysis, three things drove the swap: price (chicken ran $1.67 to $2.64 cheaper per pound at wholesale and retail in 2010), convenience products like boneless breasts and rotisserie chicken, and health concerns about saturated fat. Production got more efficient too. The average broiler weighed 5.8 pounds by 2010, up from 3.4 in 1960. Pork barely moved. 48 pounds per person in 1970, 47 in 2021. It's the most stable major protein in the American diet. Cooking oils nearly quadrupled Salad and cooking oil consumption went from 15 to 54 pounds per person. The food industry moved away from animal fats (lard and tallow) toward vegetable oils, especially soybean oil. Health guidance in the 1980s and 1990s pushed consumers toward unsaturated fats, and the growth of restaurant and processed food drove commercial demand. By 2010, cooking oils made up the largest share of added fats in the American diet. Margarine collapsed, butter held steady Margarine fell 68%, from nearly 11 pounds per person to 3.5. Research in the 1990s linked the trans fats in partially hydrogenated oils to heart disease, and consumers bailed. A USDA retrospective notes that butter overtook margarine again around 2005, despite costing nearly four times as much. The FDA's 2018 ban on partially hydrogenated oils finished the job. Butter itself held flat at 5.4 to 5.7 pounds per person. The "butter is back" story is really a "margarine left" story. Vegetables peaked in 2000 and have been sliding Total vegetable availability peaked at 425 pounds per person around 2000, then fell to 384 by 2021. The biggest single driver is potatoes. Per capita potato availability dropped 28 pounds as Americans ate fewer fries and baked potatoes. Canned and frozen vegetables also declined steadily, down 32 pounds since the mid-1990s. Fresh vegetables are the exception, rising modestly even as the total fell. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It downloaded USDA data files, verified the key stories, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: USDA Economic Research Service, Food Availability Per Capita Data System. It tracks food availability rather than direct observed intake, so it's best read as a long-run proxy for consumption. Covers 200+ commodities from 1909 to 2021 (varies by category). The cleaned dataset is available here. --- ## Coal collapsed and natural gas took over the U.S. power grid URL: https://www.randalolson.com/2026/03/26/us-electricity-generation-by-source-1950-2024/ Published: 2026-03-26 Updated: 2026-06-20 Categories: data visualization Tags: beautiful-charts-with-ai, energy, electricity, usa Coal went from generating half of U.S. electricity to just 15% in two decades. 75 years of EIA data show how fast the American power grid is changing. Part of Teaching an AI Agent to Make Beautiful Charts The U.S. generates 13 times more electricity today than it did in 1950. But the bigger story isn't how much we generate. It's where that electricity comes from. I grabbed 75 years of data from the EIA's Monthly Energy Review and charted the shifts. Coal's collapse Coal went from generating 52% of U.S. electricity in 2000 to just 15% by 2024. It peaked in 2007 at 2,016 billion kWh, then lost two-thirds of its output in less than two decades. U.S. coal capacity peaked at 317.6 GW in 2011; by the end of 2026, half of that capacity will be gone. The shale gas takeover Natural gas filled the gap. Gas tripled its output from 601 billion kWh in 2000 to 1,870 billion kWh in 2024, now generating 43% of U.S. electricity. The mechanism was price: hydraulic fracturing flooded the market with cheap shale gas, and Henry Hub spot prices crashed from $8.88/MMBtu in 2008 to $2.77 by 2012. Utilities retired coal plants and built gas turbines. Gas overtook coal in 2016. Nuclear holds steady Nuclear has hovered between 770 and 810 billion kWh annually for the past two decades, sitting at about 18% of U.S. electricity. The first commercial reactors came online in the late 1950s, and nuclear grew fast through the 1970s and 80s. Three Mile Island in 1979 and rising construction costs froze new builds, but the existing fleet keeps running. Its carbon-free output is still larger than wind and solar combined. Here's where it gets interesting: Microsoft signed a 20-year deal in 2024 to restart TMI's Unit 1 to feed its AI data centers. The plant that froze nuclear construction for a generation is being brought back online because the grid can't build new capacity fast enough. Wind and solar's breakout decade Wind was 18 billion kWh in 2005. By 2024, wind hit 452 billion kWh (10.5%) and utility-scale solar reached 220 billion kWh (5.1%), combining for 15.6% of the grid. Wind's growth owes a lot to the federal production tax credit, first enacted in 1992 when the U.S. had less than 1.5 GW of wind capacity. Solar grew 12x in a single decade, from 18 billion kWh in 2014 to 220 billion kWh in 2024. (These solar figures cover utility-scale installations only. Including rooftop solar pushes the 2024 total closer to 304 billion kWh.) Renewables as a whole (wind, solar, hydro, geothermal, biomass) generated 976 billion kWh in 2024, making up 22.7% of total generation. That's more than coal and more than nuclear. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It pulled the latest EIA data, built the chart in Python, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: U.S. Energy Information Administration, Monthly Energy Review Table 7.2a. The cleaned dataset is available here. --- ## Americans used to outlive their peers. Now they die 4 years sooner. URL: https://www.randalolson.com/2026/03/26/us-life-expectancy-vs-peer-nations-1960-2023/ Published: 2026-03-26 Categories: data visualization Tags: beautiful-charts-with-ai, public health, life expectancy, usa In 1960, Americans lived 1.5 years longer than peers in wealthy nations. By 2023, they die 4 years sooner. World Bank data shows how the U.S. fell behind. Part of Teaching an AI Agent to Make Beautiful Charts Here's a stat that should stop you cold: in 1960, the average American outlived their counterparts in Japan, the U.K., South Korea, and other wealthy nations by a year and a half. By 2023, that lead had flipped into a 4-year deficit. I pulled 64 years of World Bank life expectancy data to chart the divergence. The U.S. started ahead and fell behind In 1960, U.S. life expectancy was 69.8 years, 1.5 years above the peer average of 68.3. That lead narrowed through the 1970s and vanished by around 1980. From there, the gap reversed and kept widening. A 2023 study in the American Journal of Public Health traced the slowdown as far back as the 1960s, when U.S. gains in heart disease mortality began lagging behind Europe and Japan. The causes are structural: higher obesity rates, no universal healthcare, more gun deaths, and more car fatalities than peer nations. Our World in Data's analysis calls it a "pervasive disadvantage" spanning dozens of causes of death, not a single smoking gun. South Korea started at 54 and blew past the U.S. In 1960, South Korean life expectancy was just 53.8 years, 16 years below the U.S. By 2003, South Korea had caught up and passed the U.S. entirely. By 2023, the gap had flipped to 5 years in Korea's favor (83.4 vs. 78.4). The speed of that convergence is staggering: South Korea gained nearly 30 years of life expectancy in six decades while the U.S. gained fewer than 9. South Korea's rise tracks closely with its rapid industrialization, expansion of universal health insurance in 1989, and one of the lowest obesity rates among wealthy nations. It's a powerful counterpoint to the idea that the U.S. trajectory is somehow inevitable for rich countries. Opioids froze U.S. progress for a decade From 2010 to 2019, U.S. life expectancy barely budged, stuck between 78.5 and 78.9 years while peer nations kept climbing. The opioid epidemic is the clearest reason. Drug overdose deaths rose from 38,329 in 2010 to 105,007 in 2023, nearly tripling in just over a decade. A 2024 Lancet study estimated that opioid overdoses alone reduced U.S. life expectancy by 0.67 years in 2022. The crisis came in three waves: prescription opioids like OxyContin through the 2000s, a heroin surge starting around 2010, then synthetic fentanyl dominating after 2013. The good news: CDC data shows overdose deaths dropped 27% in 2024, down to 79,384. Still an enormous number, but the first sustained decline in years. COVID blew the gap wide open Between 2019 and 2021, U.S. life expectancy dropped 2.5 years, from 78.8 to 76.3. Peer nations lost just 0.2 years on average over the same period. Japan barely moved at all, holding steady at 84.4. A study in the American Journal of Public Health found that U.S. losses disproportionately hit young and middle-aged adults: increases in deaths among 15-to-64-year-olds accounted for nearly half the decline. In peer countries, 88% of losses came from deaths over age 65. The U.S. had roughly 1.7 million excess deaths between March 2020 and the end of 2022. By 2023, U.S. life expectancy had recovered to 78.4, and preliminary CDC data puts 2024 at 79.0. That's above the pre-pandemic 78.8, but the gap with peers remains wider than it was in 2019. Japan leads by 5.7 years, and it's not genetics In 2023, a Japanese newborn can expect to live to 84.0. An American newborn, 78.4. That's a 5.7-year gap. When Japanese people move to the U.S. and adopt American diets and driving habits, their obesity rates converge toward American levels. It's environment and behavior, not DNA. The numbers tell the story. Japan's adult obesity rate is about 4%; the U.S. rate is 40%. Japan has had universal health insurance since 1961 and mandates annual employer health screenings. A 2020 study in the European Journal of Clinical Nutrition points to diet as the key differentiator: high fish and soy intake, low red meat, minimal sugar-sweetened drinks, and widespread green tea consumption have kept Japan's heart disease and cancer mortality far below Western levels. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It pulled the latest World Bank data, built the chart in Python, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: World Bank, Life expectancy at birth (SP.DYN.LE00.IN). Peer nations in the average: Australia, Canada, France, Germany, Japan, South Korea, Sweden, and the U.K. The cleaned dataset is available here. --- ## 156 years of marriage and divorce in the United States URL: https://www.randalolson.com/2026/03/26/us-marriage-and-divorce-rates-1867-2023/ Published: 2026-03-26 Updated: 2026-06-20 Categories: data visualization Tags: beautiful-charts-with-ai, marriage, divorce, usa U.S. marriage rates have fallen 42% since the 1980s. A chart of 156 years of CDC marriage and divorce data, through the COVID-era record low. Part of Teaching an AI Agent to Make Beautiful Charts Back in 2015, I charted 144 years of U.S. marriage and divorce data after scraping dozens of CDC reports by hand. That post ended at 2014. The CDC has since published updated figures through 2023, so I brought the chart up to date. Wars and weddings The post-WWII marriage boom in 1946 hit 16.2 per 1,000 people, the highest rate in the entire dataset. Soldiers came home, and the GI Bill's zero-down-payment home loans made it possible to start a family right away. VA-backed loans funded nearly 2.4 million home purchases between 1944 and 1952. The same pattern shows up at a smaller scale around WWI: a spike when the U.S. entered the war in 1917 (11.1), a dip while the men were overseas, and another spike when they returned in 1920 (12.0). Less discussed: the only notable spike in divorce rates in 156 years of data also followed WWII, jumping to 4.3 per 1,000 in 1946. Couples who barely knew each other before one shipped overseas came home to find they were strangers. Meanwhile, women who had entered the workforce in record numbers during the war now had the economic independence to leave bad marriages. The Great Depression and economic hardship The sharpest drop in marriage rates before the modern era came during the Great Depression. By 1932, the rate had fallen to 7.9 per 1,000, a 25% decline from just a few years earlier. When Americans fell on hard times, marriage was one of the first things to take the back seat. The rate bounced back quickly once the economy recovered, reaching 10.3 by 1934. The long decline since the 1980s The marriage rate was 10.6 per 1,000 in 1980. By 2023 it had fallen to 6.1. That's a 42% decline over four decades, and unlike the war-era dips, this one shows no sign of bouncing back. Cohabitation explains some of the gap: Pew found in 2019 that among adults 18-44, more had lived with an unmarried partner (59%) than had ever been married (50%). Divorce followed a parallel arc. Rates rose sharply through the 1960s and 70s, peaked at 5.3 per 1,000 around 1979-1981, then fell steadily to 2.0 by 2023. The spike lines up with the spread of no-fault divorce laws: California passed the first one in 1969, and virtually every state followed by the early 1980s. The decline since then is probably simpler: fewer people getting married means fewer marriages to dissolve. COVID-19 set a new record low The 2020 COVID lockdowns drove marriage rates to 5.1 per 1,000, the lowest in 156 years of data. Lower than the Great Depression, lower than the Civil War era. Rates bounced back to 6.0-6.2 in 2021-2022, but 2023 sits at 6.1, right back on the declining trendline. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It pulled the latest CDC data, built the chart in Python, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: CDC National Center for Health Statistics (2000-2023), combined with historical data I scraped from CDC reports back in 2015 (1867-1999). The full merged dataset is available here. --- ## Where racial diversity lives in America, county by county URL: https://www.randalolson.com/2026/03/26/us-racial-diversity-by-county/ Published: 2026-03-26 Categories: data visualization Tags: beautiful-charts-with-ai, demographics, diversity, usa A county-level map of racial and ethnic diversity in the U.S., built from 2020 Census data. The national diversity index crossed 61% for the first time. Part of Teaching an AI Agent to Make Beautiful Charts Back in 2014, I mapped racial diversity for every U.S. county using 2010 Census data. That map got a lot of traction on Reddit, and the question people kept asking was: how does this look with newer data? The 2020 Census has been out for a while now, so I built an updated map. The map shows the Census Bureau's Diversity Index for all 3,143 counties and county equivalents in the 50 states and D.C.: the probability that two people picked at random belong to different racial or ethnic groups. Zero means everyone is the same race. The national index hit 61% in 2020, up from 55% in 2010. But the median county sits at just 33%. Most Americans live in diverse places; most places in America are not diverse. If you want the before-and-after view, I also mapped where diversity changed fastest from 2010 to 2020. California metros lead the mainland in diversity The darkest clusters on the mainland are California's major metro areas. Alameda County (Oakland) scores 75%, Sacramento 73%, San Francisco 70%. What these counties have in common is a genuine four-way mix: large White, Hispanic, Asian, and Black populations all in one place. Hawaii tops the country at 76% statewide, driven by its mix of Asian, White, Native Hawaiian, and multiracial residents. Texas has one county in the same league: Fort Bend, outside Houston, scores 76%, the highest in the state. It works for the same reason California's metros do. No single group dominates, and large White, Hispanic, Black, and Asian populations all live in the same county. The NYC metro is the most diverse region in the East Queens County scores 77% on the diversity index, the most diverse large county in the eastern United States. Brooklyn hits 75%. The dark patch extends across northern New Jersey (Essex, Hudson, and Middlesex counties all above 70%) and into the Maryland suburbs of D.C. (Montgomery County at 74%). This is where immigration from dozens of countries concentrates in dense, mixed neighborhoods. Look at the map and you can see the pattern: every major metro area pops against its surroundings. The Twin Cities stand out against a pale Minnesota. Denver against Colorado. Atlanta against Georgia. Immigration, jobs, and universities all pull people to the same places. The rural Midwest and Appalachia form the broadest low-diversity region The lightest band on the map stretches from West Virginia through Kentucky, across Iowa and Indiana, and up into the Dakotas. Kentucky has 51 counties below 15% on the diversity index. West Virginia has 36 out of 55. Iowa has 46 out of 99. These are places that largely missed both big immigration waves and the main destinations of the Great Migration. The exceptions pop on the map: meatpacking towns in Iowa and Nebraska that attracted Hispanic workers, and university counties with international students. But away from the coasts and major metros, the center of the country remains overwhelmingly White. Florida metros show how far the pattern has spread Florida gives the South another dark patch on the map, but the bigger story is how common highly diverse metro counties have become outside the old coastal gateways. Broward County scores 72%, the highest in Florida. Orange County, home to Orlando, is right behind at 71%. Fort Bend outside Houston hits 76%, Gwinnett outside Atlanta 75%, and Prince William outside Washington 74%. Brookings found the same pattern in the 2020 Census: America's big suburbs are now often more diverse than the country as a whole. Miami-Dade shows why the details matter. It feels like the kind of place that should top the chart, but the county's population is less evenly split than Broward's or Orange's. Hispanics make up 69% of residents there, so its diversity index lands at 49%. The index rewards balance across big groups, not just a place drawing people from everywhere. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It pulled 2020 Census data via the Census Bureau API, computed the Gini-Simpson diversity index for all 3,143 counties and county equivalents in the 50 states and D.C., and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: U.S. Census Bureau, 2020 Decennial Census, Table P2 (Hispanic or Latino, and Not Hispanic or Latino by Race). The cleaned dataset with diversity index scores for all counties is available here. --- ## 96% of U.S. counties grew more diverse from 2010 to 2020 URL: https://www.randalolson.com/2026/03/26/us-racial-diversity-change-2010-2020/ Published: 2026-03-26 Categories: data visualization Tags: beautiful-charts-with-ai, demographics, diversity, usa A county-level map of how racial diversity changed across America between the 2010 and 2020 Census. The Northeast and Midwest diversified fastest. Part of Teaching an AI Agent to Make Beautiful Charts In the companion post, I mapped where racial diversity lives in America today. This post answers the follow-up question: where is it changing fastest? I computed the diversity index for every matched county and county equivalent in the 50 states and D.C. in both the 2010 and 2020 Census, then mapped the difference. Orange means a county grew more diverse; blue means less diverse. The map is almost entirely orange. Out of 3,139 matched counties and county equivalents, 3,020 (96.2%) grew more diverse over the decade, gaining 5.9 points on average. Only 119 moved in the other direction. The Northeast posted the largest average gains Northeast counties gained an average of 7.9 points on the diversity index, with 99.5% of them trending more diverse. No other region was this consistently orange. Suburban counties outside New York, Philadelphia, and Boston absorbed immigration from Asia and Latin America that had previously concentrated in the cities. Luzerne County, Pennsylvania (Wilkes-Barre) gained 17 points. Its Hispanic population went from about 21,000 residents in 2010 to about 47,000 in 2020. Similar shifts played out across dozens of mid-sized counties in Pennsylvania, New Jersey, and Connecticut. North Dakota oil counties had the biggest single-county gains Williams County, North Dakota (Williston) gained 25 points on the diversity index, the single biggest increase in the country. The Bakken oil boom drew workers from across the country, nearly doubling the county's population from 22,000 to 41,000 in a single decade. Neighboring Stark County gained 17 points for the same reason. The Midwest as a whole gained 6.9 points on average, with 98.9% of counties growing more diverse. Beyond the oil patch, Hispanic immigration to meatpacking towns was the clearest driver. Pew Research found that 214 Midwest counties had Hispanic population growth of 50% or more between 2010 and 2020, concentrated around food processing and agriculture jobs. Southeast suburbs around Atlanta, Charlotte, and Nashville surged Forsyth County, Georgia, north of Atlanta, gained 21 points on the diversity index. Its Asian population more than quadrupled, from about 11,000 residents to 45,000, and reached 18% of the county by 2020. The county as a whole grew 43% in a decade. Cabarrus County outside Charlotte and Rutherford County outside Nashville each gained 14 points. The pattern is simple: where jobs and housing are being built, people from different backgrounds move in. The South overall gained a more modest 4.9 points on average, dragged down by slow-changing rural counties in Appalachia and the Deep South. The blue counties are the exception, not the story The blue counties do not add up to a separate national trend. There are only 119 of them, and most are places where one group already dominated and then pulled a little farther ahead. Texas had the most, 39 out of 254 counties. Starr County went from 95.7% Hispanic in 2010 to 97.7% in 2020, and Imperial County in California went from 80% Hispanic to 85%. On this map, "less diverse" usually just means the local mix became a little less even. Why did this happen nationwide? Three things happened at once. The non-Hispanic White population declined by 5.1 million, the first such drop in U.S. history, as deaths outpaced births in an aging population. Meanwhile, the Hispanic population grew by 11.2 million and the Asian population by 5.2 million, together accounting for nearly all national growth. And the multiracial population jumped from 9 million to 33.8 million, though much of that reflects changes in how the Census Bureau processed responses rather than actual demographic shifts. Add it up, and in nearly every county in America, the direction is the same. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It pulled both 2010 and 2020 Census data via the Census Bureau API, computed the diversity index change for all 3,139 matched counties and county equivalents in the 50 states and D.C., and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data source: U.S. Census Bureau, 2010 & 2020 Decennial Census, Table P2 (Hispanic or Latino, and Not Hispanic or Latino by Race). The cleaned dataset with diversity index scores and change values for all counties is available here. --- ## Working your way through college now takes 5x more hours than in 1970 URL: https://www.randalolson.com/2026/03/26/working-your-way-through-college-1963-2024/ Published: 2026-03-26 Updated: 2026-06-20 Categories: data visualization Tags: beautiful-charts-with-ai, education, college, minimum wage, usa In 1970, a student could work 5 hours per week at minimum wage to cover public university tuition. By 2022, that number hit 26. Part of Teaching an AI Agent to Make Beautiful Charts In 2014, a rant about Michigan State University's tuition went viral and people pushed back, saying it was just one school. So I followed up with a proper analysis using national data. That post used IPEDS tuition data from 1987-2010 and found the same trend: a 1979 student could pay tuition with a part-time summer job (182 hours), while a 2013 student needed a full-time job for half the year (991 hours, extrapolated). Over 5x as many hours for the same education. Twelve years later, the NCES data now goes back to 1963 and forward to 2023, and I added private universities to the picture. I pulled tuition data from the NCES Digest of Education Statistics, combined it with federal minimum wage history, and calculated hours per week a student would need to work year-round to cover annual tuition and fees. The actual 2013 data, by the way, came in at 1,147 total hours, worse than the 991 my model had predicted. A summer job used to cover tuition Through the 1970s and into the early 1980s, a student at a public university could cover annual tuition with about 5 hours of minimum wage work per week. That's roughly 260 hours total for the year. A full-time summer job at minimum wage paid about 520 hours. Tuition was covered with room to spare. This was the era people mean when they say "I worked my way through college." They're not wrong about their own experience. Public university tuition in 1970-71 was $394. The federal minimum wage was $1.60/hour. The math worked. Two forces broke the math Look at the public line in the early 1980s. It starts climbing right when Congress froze the federal minimum wage at $3.35 in 1981, where it stayed for nine years. Tuition kept rising while the wage floor sat still, and the hours needed nearly doubled from 5 to 10 by 1990. Sound familiar? The same thing is happening right now with the post-2009 freeze. Then states started pulling back their funding. From 2000 to 2015, state funding per student fell roughly 30% after adjusting for inflation, according to Pew. The Center on Budget and Policy Priorities found that published tuition at four-year public colleges rose 35% in the years following the 2008 recession, as states slashed budgets and never fully restored the funding. Universities passed the shortfall directly to students. Private tuition crossed the full-time work line in 1988 In the 1987-88 academic year, covering private university tuition first required more than 40 hours of minimum wage work per week, year-round. More than a full-time job, just for tuition. No food, no rent, no books. That was nearly 40 years ago, and it has only gotten further out of reach since. By 2022-23, private nonprofit tuition required 102 hours of minimum wage work per week. There are only 168 hours in a week. Working your way through private school stopped being difficult a long time ago. It is now physically impossible. Minimum wage increases help, then the freeze makes it worse The chart shows something interesting around 2007-2009. Congress raised the minimum wage in three steps, from $5.15 to $7.25, and the hours needed to cover private tuition dropped from 82 to 68 per week. Minimum wage increases actually work as temporary relief. Then the freeze started, and by 2022 the line had climbed past 100. The federal minimum wage has been $7.25 since July 24, 2009. That's the longest period without an increase since the minimum wage was established in 1938. During that same period, average public university tuition rose from $6,717 to $9,750, a 45% increase, while the wage stayed frozen. Thirty states and D.C. have set their own minimum wages above $7.25, but students in the remaining 20 states are stuck with a rate that hasn't moved in 17 years. Where things stand now As of 2022-23, a student working at the federal minimum wage needs to work 26 hours per week, year-round, just to cover public in-state tuition and fees. That's before paying for housing, food, textbooks, or transportation. At a private nonprofit university, the number is 102 hours per week. The College Board's 2024-25 figures are even higher: $11,610 for public and $43,350 for private, which would push those to 31 and 115 hours per week respectively. The next time someone tells a college student to "just work harder," show them this chart. The math that worked in 1975 stopped working decades ago. How this chart was made An AI agent built this chart end-to-end as part of the Beautiful Charts with AI series. It pulled tuition data from NCES, combined it with federal minimum wage history, and iterated on the design until it passed the Tufte Test, a data visualization quality standard from Goodeye Labs. The workflow behind it is public: run the same high-signal chart workflow to make your own. Data sources: NCES Digest of Education Statistics, Table 330.10 (tuition and required fees, 1963-64 through 2022-23) and U.S. Department of Labor (federal minimum wage history). The full dataset is available here. --- ## The "Are You Sure?" Problem: Why Your AI Keeps Changing Its Mind URL: https://www.randalolson.com/2026/02/07/the-are-you-sure-problem-why-your-ai-keeps-changing-its-mind/ Published: 2026-02-07 Categories: ai, production Tags: ai, sycophancy, llm, decision-making, risk, reliability Ask your AI 'are you sure?' and watch it flip. Models fold 60% of the time because we trained them to please, not push back. The fix isn't better prompts. Try this experiment. Open ChatGPT, Claude, or Gemini and ask a complex question. Something with real nuance, like whether you should take a new job offer or stay where you are, or whether it's worth refinancing your mortgage right now. You'll get a confident, well-reasoned answer. Now type: "Are you sure?" Watch it flip. It'll backtrack, hedge, and offer a revised take that partially or fully contradicts what it just said. Ask "are you sure?" again. It flips back. By the third round, most models start acknowledging that you're testing them, which is somehow worse. They know what's happening and still can't hold their ground. This isn't a quirky bug. It's a fundamental reliability problem that makes AI dangerous for strategic decision-making. AI Sycophancy: The Industry's Open Secret Researchers call this behavior "sycophancy," and it's one of the most well-documented failure modes in modern AI. Anthropic published foundational work on the problem in 2023, showing that models trained with human feedback systematically prefer agreeable responses over truthful ones. Since then, the evidence has only gotten stronger. A 2025 study by Fanous et al. tested GPT-4o, Claude Sonnet, and Gemini 1.5 Pro across math and medical domains. The results: these systems changed their answers nearly 60% of the time when challenged by users. These aren't edge cases. This is default behavior, measured systematically, across the models millions of people use every day. Answer Flip Rate When Users Challenge AI Answer Flip Rate When Users Challenge AI Source: Fanous et al. 2025 ~58% GPT-4o ~56% Claude Sonnet ~61% Gemini 1.5 Pro All major models flip answers over half the time when challenged And in April 2025, the problem went mainstream when OpenAI had to roll back a GPT-4o update after users noticed the model had become excessively flattering and agreeable. Sam Altman publicly acknowledged the issue. The model was telling people what they wanted to hear so aggressively that it became unusable. They shipped a fix, but the underlying dynamic hasn't changed. Even when these systems have access to correct information from company knowledge bases or web search results, they'll still defer to user pressure over their own evidence. The problem isn't a knowledge gap. It's a behavior gap. We Trained AI to Be People-Pleasers Here's why this happens. Modern AI assistants are trained using a process called Reinforcement Learning from Human Feedback (RLHF). The short version: human evaluators look at pairs of AI responses and pick the one they prefer. The model learns to produce responses that get picked more often. The problem is that humans consistently rate agreeable responses higher than accurate ones. Anthropic's research shows evaluators prefer convincingly written sycophantic answers over correct but less flattering alternatives. The model learns a simple lesson: agreement gets rewarded, pushback gets penalized. How RLHF Creates a Sycophancy Loop How RLHF Creates a Sycophancy Loop The training process rewards agreement over accuracy AI Generates Two Responses Human Evaluator Picks Preferred One Agreeable Answer Wins More Often (even if less accurate) Model Learns: "Agreement = Reward" Future Responses Prioritize Validation Over Accuracy Cycle Repeats Result: Models get better at telling you what you want to hear The longer you talk with AI, the more it agrees with you. This creates a perverse optimization loop. High user ratings come from validation, not accuracy. The model gets better at telling you what you want to hear, and the training process rewards it for doing so. It gets worse over time, too. Research on multi-turn sycophancy shows that extended interactions amplify sycophantic behavior. The longer you talk with these systems, the more they mirror your perspective. First-person framing ("I believe...") significantly increases sycophancy rates compared to third-person framing. The models are literally tuned to agree with you specifically. Can this be fixed at the model layer? Partially. Researchers are exploring techniques like Constitutional AI, direct preference optimization, and third-person prompting that can reduce sycophancy by up to 63% in some settings. But the fundamental training incentive structure keeps pulling toward agreement. Model-level fixes alone aren't sufficient because the optimization pressure that creates the problem is baked into how we build these systems. The Strategic Risk You're Not Measuring For simple factual lookups, sycophancy is annoying but manageable. For complex strategic decisions, it's a real risk. Consider where companies are actually deploying AI. A Riskonnect survey of 200+ risk professionals found that the top uses of AI are risk forecasting (30%), risk assessment (29%), and scenario planning (27%). These are exactly the domains where you need your tools to push back on flawed assumptions, surface inconvenient data, and hold a position under pressure. Instead, we have systems that fold the moment a user expresses disagreement. The downstream effects compound quickly. When AI validates a flawed risk assessment, it doesn't just give a bad answer. It creates false confidence. Decision-makers who would have sought a second opinion now move forward with unearned certainty. Bias gets amplified through decision chains. Human judgment atrophies as people learn to lean on tools that feel authoritative but aren't reliable. And when something goes wrong, there's no accountability trail showing why the system endorsed a bad call. Brookings has written about exactly this dynamic in their analysis of how sycophancy undermines productivity and decision-making. To be clear: this is about complex, judgment-heavy questions. AI is plenty reliable for straightforward tasks. But the more nuanced and consequential the decision, the more sycophancy becomes a liability. Give AI Something to Stand On The RLHF training explains the general tendency, but there's a deeper reason the model folds on your specific decisions: it doesn't know how you think. It doesn't have your decision framework, your domain knowledge, nor your values. It fills those gaps with generic assumptions and produces a plausible answer with zero conviction behind it. That's why "are you sure?" works so well. The model can't tell if you caught a genuine error or you're just testing its resolve. It doesn't know your tradeoffs, your constraints, or what you've already considered. So it defers. Sycophancy isn't just a training artifact. It's amplified by a context vacuum. The Context Vacuum: Why AI Folds Under Pressure The Context Vacuum What happens when the model doesn't know how you make decisions Without Your Context AI Model Decision framework generic Domain knowledge generic Values & priorities generic "Are you sure?" Folds. Changes answer. With Your Context Embedded AI Model Decision framework yours Domain knowledge yours Values & priorities yours "Are you sure?" Holds ground. Asks for more. What you need is for the model to push back when it doesn't have enough context. It won't unless you tell it to. Here's the irony: once you instruct it to challenge your assumptions and refuse to answer without sufficient context, it will, because pushing back becomes what you asked for. The same sycophantic tendency becomes your leverage. Then go further. Embed your decision framework, domain knowledge, and values so the model has something real to reason against and defend. Not through better one-off prompts, but through systematic context that persists across how you work with it. This is the real fix for sycophancy. Not catching bad outputs after the fact, but giving the model enough information about how you make decisions that it has something to stand on. When it knows your risk tolerance, constraints, and priorities, it can tell the difference between a valid objection and pressure. Without that, every challenge looks the same, and agreement wins by default. Try It Yourself Try the experiment from the opening. Ask your AI a complex question in your domain. Challenge it with "are you sure?" and watch what happens. Then ask yourself: have you given it any reason to hold its ground? The sycophancy problem is known, measured, and model improvements alone won't fix it. The question isn't whether your AI will fold under pressure. The research says it will. The question is whether you've given it something worth defending. --- ## Why Custom Evals Matter for Production LLMs URL: https://www.randalolson.com/2025/12/22/why-custom-evals-matter-for-production-llms/ Published: 2025-12-22 Categories: ai, production Tags: ai, production, evaluation, llm Generic benchmarks measure breadth across many tasks. Your application needs depth on one specific task. This is why custom evals matter for production LLMs. If you've shipped an LLM project into production before, here's a scenario that might sound familiar. One of the major AI labs ships a new flagship model. The release notes promise improvements across the board. The benchmark results look better. You run a few test queries with the new model in your project and the responses seem good. So you make the switch and push to prod. Then a week later, you're debugging why your production application is behaving worse than before. Users are complaining. Bug reports are piling up. The model scores higher on benchmarks, but it's performing worse on what you actually need it to do. This happens because generic benchmarks measure breadth across many tasks. Your application needs depth on one specific task. The only way to know if a model will work for you is to evaluate it on your data. This is why custom evals matter for production LLMs. Let me show you what that looked like on a speech recognition project where custom evals made all the difference. Generic benchmarks show improvement, but production applications can still perform worse on your specific use case. The Problem We Faced I worked on an education technology project for early elementary students learning to speak phonemes and letter names correctly. Think of it as helping kids master the fundamental building blocks of speaking and reading. We built the initial app using OpenAI's audio model. It worked reasonably well, but we had a major blind spot. We had no idea where it was actually failing. Some sounds were consistently misclassified, but we didn't know which ones or why. Without that visibility, we couldn't improve the product or make informed decisions about switching models. The solution wasn't complicated in theory, but it required real work. We collected over 500 audio traces from actual students speaking in real classrooms. These weren't clean studio recordings. This was chaotic classroom audio with background noise, kids talking over each other, and all the messiness of the real world. We labeled each recording with whether the model's classification was correct or incorrect, plus what the actual sound was supposed to be. Real classroom conditions with background noise and chaos. This is the messy reality where our model needed to perform. That gave us a domain-specific eval set. Not a generic "how good is this model at speech recognition" benchmark, but a specific answer to "how good is this model at recognizing the phonemes that matter for early readers in real classroom conditions." This same approach works for any domain-specific LLM application. Document extraction, customer support, code generation, whatever. The principle is the same. Evaluate on your actual use case with your actual data. What the Custom Eval Revealed Avoided a Costly Migration When a new OpenAI audio model came out with claims of improvements, we were ready to jump on it. Then we ran it through our custom eval. The results were all over the place. Better on some phonemes, worse on others. The overall improvement wasn't worth the migration cost and risk of introducing new failure modes we'd have to debug. This is the trap of trusting general improvement claims. A model can be better in some ways but perform worse on what you actually need. Whether you're evaluating GPT-5 versus Claude Opus or comparing embedding models, custom evals tell you what generic claims can't. Found Where the Real Problems Were Our eval revealed specific patterns. Nasal sounds like /m/ and /n/ were consistently confused because they sound acoustically similar. Plosive consonants like /k/ and /p/ got filtered out by noise-reducing classroom microphones that thought they were just background noise. Knowing this let us focus our improvement efforts. We weren't randomly trying different prompts or model settings hoping something would work better. We knew exactly which phonemes needed attention and could measure whether our changes actually helped. For the plosive issue, the fix wasn't changing the model at all. It was upgrading to better headsets for the students. Every domain has these edge cases that generic testing misses. The only way to find them is to look at your own data. Shipped Updates Faster Once we had the eval in place, we could quickly test new approaches. Trying a different prompt structure? Run it through the eval. Adjusting parameters? Test it against the eval. We went from "try it and see how it feels" to quantitative comparison. That changed our iteration cycle from days to hours. No more guessing whether a change actually made things better or just felt better in the moment. Deployed with Confidence The eval gave us a clear bar for quality. We knew when accuracy was good enough for real students using the app. We could monitor whether performance degraded over time as OpenAI updated their models or we made changes to our pipeline. We weren't flying blind anymore, which matters when you're building something that affects how kids learn to read. Why This Matters This pattern applies to any production LLM application. Generic improvement claims are built to make models sound better broadly. But production needs specificity. You need to know if a model works for your particular use case with your particular data in your particular conditions. Building domain-specific evals takes time. For us, it meant collecting 500+ audio samples, labeling them correctly, and setting up the infrastructure to run new models and prompts through the eval consistently. That's real work. But it pays for itself immediately. We avoided a bad migration that would have cost us debugging time and potentially degraded the product. We found and fixed the actual problems instead of guessing. We shipped improvements faster because we could measure them objectively. You can't know if a model works for your use case without evaluating it on your data. The improvement claims won't tell you. Only your own eval will. Every production application needs this kind of evaluation infrastructure. Building this infrastructure is real work, which is why we're focused on making it easier at Goodeye Labs. If you're facing similar challenges figuring out whether your LLM actually works the way you need it to, check out what we're building or get in touch. --- ## How to Choose the Right LLM for AI-Assisted Coding URL: https://www.randalolson.com/2025/11/25/how-to-choose-right-llm-ai-assisted-coding/ Published: 2025-11-25 Categories: programming, ai Tags: programming, ai, productivity, workflow Benchmark scores don't predict real-world AI coding performance. Here's how to actually choose the right model for your workflow. New models drop every week, each claiming to top the coding benchmarks. How do you actually pick the right one for AI-assisted coding? The short answer: ignore the hype and focus on what actually predicts performance in your workflow. The Problem with Coding Benchmarks Traditional coding benchmarks like HumanEval measure isolated function-writing tasks, and even more sophisticated benchmarks like SWE-bench test single-shot issue resolution. A model can ace these benchmarks and still struggle in real agentic coding workflows, which require iterative multi-turn collaboration, effective tool use, and managing context across files over extended sessions. The Artificial Analysis Coding Index ranks models on traditional coding benchmarks. But coding benchmarks don't tell the whole story. Source: artificialanalysis.ai, November 2025. The Agentic Index: A Better Signal I've found the Artificial Analysis Agentic Index to be a much better predictor of how a model will perform for AI-assisted coding. It measures how well models handle tool use and multi-step problem solving. For coding, this translates to how well a model navigates codebases, coordinates changes across files, uses search and editing tools, and maintains context over extended sessions. A good example: DeepSeek R1 0528 scores well on coding benchmarks but poorly on the agentic index. It might be fine for writing standalone scripts, but it won't make effective use of the tools in modern AI coding platforms. In practice, this means it might write good code when given explicit instructions, but may struggle to navigate your codebase, search for relevant context, or use tools and MCPs effectively. The rankings shift significantly when you measure what actually matters. The Agentic Index measures tool use and multi-step reasoning. Notice how some models drop significantly from their coding benchmark rankings. Source: artificialanalysis.ai, November 2025. Evaluating Trade-offs You don't always need the absolute best model. Artificial Analysis provides scatter plots that let you visualize trade-offs between performance, cost, and speed. Look for models in the upper-right quadrant: high performance, lower cost. The sweet spot is often models one tier below the frontier that deliver 90% of the capability at a fraction of the price. Performance vs. price trade-offs. The green quadrant shows models with strong performance at lower cost. Source: artificialanalysis.ai, November 2025. Performance vs. speed trade-offs. For implementation work, speed often matters more than raw intelligence. Source: artificialanalysis.ai, November 2025. Match the Model to the Phase In my previous post, I covered the three phases of AI-assisted coding: planning, implementation plan creation, and implementation. Different phases have different requirements. For planning and creating the implementation plan, I use high-reasoning models. Mistakes at this stage are expensive: a flawed plan means wasted implementation time, or worse, shipping bugs that a more thorough analysis would have caught. I want thorough analysis, not speed. For implementation, I switch to faster, cheaper models. The thinking is already done; the AI just needs to execute a clear plan. Right now, I use Claude Opus 4.5 for planning and implementation plan creation, then switch to Cursor's Composer 1 model for implementation. These specific models change constantly as new ones release, so use artificialanalysis.ai to stay current on the trade-offs. Always Vibe Check Even the best benchmarks can mislead. Gemini 3 Pro currently dominates both the coding and agentic benchmarks, but when I tried it for planning, it kept jumping straight to solutions instead of exploring the problem space with me. It failed the vibe check for that use case, despite topping the leaderboards. Always test models in your actual workflow before committing. Run 2-3 representative tasks from your actual workflow. Pay attention to whether the model collaborates or steamrolls, whether it asks clarifying questions, and whether it uses tools appropriately. Benchmarks are a starting point, not the final answer. The Takeaway Don't chase benchmark leaderboards blindly. Use the agentic index as a better starting point, evaluate trade-offs based on your phase and priorities, and vibe check before committing. Artificial Analysis is your friend for staying current as the model landscape shifts. I teach workshops on AI-assisted development techniques, helping teams build effective workflows with these tools. If you're interested in leveling up how your organization uses AI for coding, let's talk. --- ## The Three Phases of AI-Assisted Coding URL: https://www.randalolson.com/2025/11/24/three-phases-ai-assisted-coding/ Published: 2025-11-24 Categories: programming, ai Tags: programming, ai, productivity, workflow Most developers skip straight to asking AI to write code. Here's the three-phase workflow that makes AI-assisted coding actually work. Most developers approach AI coding assistants backwards. They open their IDE, type "write me a function that does X," and expect magic. When the AI produces something that doesn't quite fit, they blame the tool. "AI coding is overhyped," they say after a frustrating week of wrestling with generated code that keeps missing the mark. But here's what I've learned after using AI coding tools daily since 2023: the problem isn't the AI. It's that most people skip straight to implementation without doing the crucial work that makes implementation trivial. The insight is simple: separate thinking from doing. AI-assisted coding works best when you treat it as three distinct phases, not one. And if you're doing it right, 75% of your time should be spent in the first two phases. The actual coding should be the easy part. The three phases of AI-assisted coding: most of your time should be spent in the first two phases. The Three Phases Phase 1: Interactive Planning This is where most people skip ahead, and it's exactly where you should slow down. In the planning phase, you're not asking the AI to write code. You're having a conversation. You're exploring the problem space together, considering different approaches, and making architectural decisions. The AI becomes your thought partner, not your code generator. I explicitly ask the AI to ask me high-information questions. Instead of the AI just proposing solutions, I have it research the codebase and relevant sources, then ask me questions that will shape the direction. I provide answers, it researches more, asks more questions, and we iterate until we've thoroughly explored the problem. This back-and-forth is where the real value happens. You're injecting your expertise and project context into the planning process. The AI brings broad knowledge and can spot things you might miss. Together, you arrive at a much better plan than either would alone. This kind of deep, exploratory work benefits from high-reasoning models. Mistakes at this stage are expensive. If you have to restart mid-implementation because you missed something, or worse, you ship a bug because you didn't think through an edge case, you'll wish you'd spent more time here. Here's the prompt I use: I want to plan [describe the feature or change you're building]. Before we start implementing anything, I want you to thoroughly research the codebase, any relevant documentation, and best practices from reliable sources on the web. Then ask me high-information questions to help decide the best approach and how to prioritize the work. For each question: 1. Ask the question 2. Provide your reasoning on how you would answer it 3. Give your recommendation Focus on strategic questions, not implementation details. We're planning, not coding yet. When you have no more high-information questions, tell me and wait for my next instruction. A note on scope: When I'm iterating rapidly toward an MVP or prototype, I add a line to this prompt telling the AI to keep things simple and focus on high-ROI changes. This keeps the AI grounded when you just need something that works. If you're building enterprise software with long-term maintenance in mind, you might skip this constraint. The magic is in that last line of the prompt. The AI signals when it's done exploring, and you decide when to move to the next phase. You stay in control of the process. Phase 2: Implementation Plan Creation Once you've explored the problem and made your key decisions, it's time to create a detailed implementation plan. This is still not coding. This is creating the roadmap that makes coding straightforward. Write the implementation plan assuming the agent that implements it will have zero context from your planning conversation. This might seem redundant, but it's essential. You're going to hand this plan to a fresh AI session (or a future version of yourself), and it needs to stand alone. I have the AI write the plan to a markdown file in my repo with: The current state of the relevant code and any important context upfront Step-by-step tasks that can be checked off Enough detail that each step is unambiguous Checkpoints after major features where you pause to test, lint, and git commit your changes before moving on Like Phase 1, this still benefits from high-reasoning models. You're still thinking, not doing. Here's the prompt I use: Based on our planning discussion, create a detailed implementation plan. Write it to a markdown file at [path]. The plan should: - Start with the current state of the relevant code and any important context - Break the work into clear, sequential steps - Include enough detail that someone with no context from this conversation could execute it - Add checkpoints after major features where we should pause to test, verify everything works, and git commit before continuing - Use checkbox format so progress can be tracked Be thorough but don't write the actual code. The goal is a roadmap, not implementation. The checkbox format matters. As you implement, you (or the AI) check things off and can add notes. It becomes a living document that tracks progress. Phase 3: Implementation Now, finally, you write code. But if you've done Phases 1 and 2 well, this phase should feel almost effortless. The hard decisions are made. The plan is clear. You're just executing. Start a fresh chat session. Point the AI at your implementation plan file and nothing else. Then tell it to follow the plan. Read the implementation plan at [path]. We're going to work through it step by step. Before each change, study the relevant code thoroughly. Make minimal, precise changes. After completing each step, mark it done in the plan. Start with step 1. Because the plan is comprehensive and context-free, the AI doesn't need to understand the full history of your planning discussions. It just needs to follow clear instructions. And when everything clicks, you can sit back and watch as the AI works through the plan, checking off steps as it goes. Before you know it, you're done. That moment never stops feeling a little bit magical. Since the thinking is already done, this phase works fine with lighter, faster models. You're optimizing for speed and cost now, not deep reasoning. The 75/25 Rule If you're spending most of your time in Phase 3, you're doing it wrong. Phases 1 and 2 (planning and plan creation) should take roughly 75% of your time on a feature. Phase 3 (implementation) should be about 25%. When implementation feels hard or the AI keeps producing code that doesn't fit, that's a signal your plan wasn't detailed enough. Go back and improve it. This feels counterintuitive. We want to see code happening. Planning feels like we're not making progress. But the time you invest in planning pays off many times over in smoother implementation and fewer rewrites. The Takeaway AI coding assistants are powerful, but they're not magic code generators. They're thought partners that become code generators. The developers who get the most value from these tools are the ones who lean into the partnership, who use the AI to think through problems before asking it to write solutions. Separate thinking from doing. Plan before you implement. And when implementation feels easy, you'll know you did the planning right. I teach workshops on AI-assisted development techniques, helping teams build effective workflows with these tools. If you're interested in leveling up how your organization uses AI for coding, let's talk. --- ## 3 AI Solutions That Actually Drive Business Value in 2025 URL: https://www.randalolson.com/2025/06/13/3-ai-solutions-business-value-2025/ Published: 2025-06-13 Categories: technology, ai, business Tags: ai, technology, business strategy Three proven AI solutions that consistently deliver measurable business value in 2025, with practical implementation strategies and real-world examples. Last time, I wrote about how AI has moved from "experimental curiosity" to "must-have business tool" in 2025. The question I keep getting from clients isn't "Should we use AI?" anymore. It's "What AI solutions actually work?" After architecting AI solutions for myriad companies this year, I've identified several distinct patterns that consistently deliver results. While there are many approaches that work, I'm going to focus on three specific solutions that stand out as the clear winners. These aren't moonshot ideas or bleeding-edge experiments. They're practical solutions that consistently deliver measurable ROI across different industries and company sizes. The Company Brain: Knowledge Base + AI Agent The Problem: Your company's most valuable information is scattered everywhere. Slack channels, SharePoint folders, random Google Docs, institutional knowledge walking out the door when employees leave. Finding the right information takes forever, and synthesizing insights across departments is nearly impossible. The Solution: A RAG (Retrieval-Augmented Generation) system that ingests all your company data and pairs it with an AI agent that can actually understand and synthesize information rather than just search for keywords. This isn't just a fancy search engine. The AI can connect dots across documents, maintain your company's voice and tone, and provide context-aware answers. Need to onboard a new hire? The AI can pull relevant policies, team structures, and project histories. Writing a client proposal? It can reference past successful proposals and current company capabilities. Getting Started: You can have a basic version running in 1-2 hours using ChatGPT Custom GPTs or Claude Projects. Upload your key documents, set some context about your company voice, and you've got an internal knowledge assistant. Scale up to dedicated platforms like Pinecone when you need enterprise security or larger document volumes. The AI Thought Partner The Problem: Knowledge work is increasingly complex, but most of us are still thinking in isolation. We need someone who knows both the business context AND our personal work patterns to help us process ideas, check our thinking, and spot blind spots. The Solution: This is the evolution of the company knowledge base. An AI system that knows your company data but also learns your decision-making style, work patterns, and strategic thinking preferences. It's designed to work alongside you on specific tasks, from strategic brainstorming to financial analysis to project planning. The key difference is personalization. This AI learns that you prefer data-driven arguments, tend to be overly conservative in financial projections, or always forget to consider regulatory implications. It adapts its recommendations accordingly. Getting Started: Enable memory features in ChatGPT or Claude and consistently feed it context about your role, decision-making preferences, and company dynamics. Perplexity Pro works great for research-heavy thought partnership. The AI gets smarter about your needs over time. Smart Automation Agents The Problem: Traditional automation breaks down when processes involve variations, exceptions, or judgment calls. But there's a huge category of work that's "dynamic but repetitive." Tasks that require understanding context but follow predictable patterns. The Solution: AI agents that can handle process automation requiring cognitive capabilities like interpreting meaning, recognizing patterns, and making judgment calls that rule-based systems can't handle. Here are two examples I see delivering consistent high ROI: Document Processing & Data Extraction Processing insurance claims, loan applications, or vendor contracts involves understanding unstructured documents with varying formats, handwriting, and context. Traditional automation chokes on format variations and can't extract meaning from complex documents. A global insurance company implemented an AI-driven document processing solution and reduced claim processing time by 80% while improving data accuracy to 99%. A leading healthcare provider digitized their patient records management, allowing them to extract critical data from handwritten notes and forms, resulting in a 60% reduction in manual data entry and a 40% decrease in processing costs. A large retail chain automated their invoice processing workflow, enabling them to process invoices 70% faster, reduce late payments by 90%, and reallocate staff to higher-value tasks, ultimately saving millions in annual operational costs. Intelligent Lead Qualification & Routing Determining which sales inquiries deserve immediate attention requires analyzing behavioral patterns, communication tone, company fit, and timing signals. Rule-based systems can't interpret intent from unstructured inquiries or predict conversion likelihood from complex data patterns. Real estate companies implementing AI-powered lead qualification are seeing significant improvements in conversion rates through better matching of leads to appropriate agents based on specialties, availability, and geographic focus. The AI handles initial qualification calls, gathers key information, and automatically routes leads to the best-matched human agents, eliminating manual sorting and reducing time to first contact. Getting Started: Zapier handles simple AI-powered automations between popular apps. Perfect for basic document processing and lead routing workflows. n8n works better for complex logic or custom integrations. Start with your most painful repetitive process and automate one piece at a time. Implementation Reality Check Pick the solution that addresses your biggest pain point, not the one that seems simplest. Document automation might be your easiest win if you're drowning in manual processing, while a knowledge base could be more complex than a basic thought partner setup. The key insight: all three solutions can start small and deliver immediate value. A basic ChatGPT Custom GPT handling support questions beats a perfect enterprise system that never launches. Focus on solving one real problem well, then expand from there. In 2025, it's no longer acceptable to use AI that isn't grounded in current, relevant data, whether that's live web search or your company's proprietary information. These three solutions give you practical ways to make that happen without boiling the ocean. The companies winning with AI aren't the ones with the fanciest tech stack. They're the ones that picked the right problem, started simple, and iterated based on real results. --- ## Key AI Technologies That Shaped 2024 and Are Driving Business Value in 2025 URL: https://www.randalolson.com/2025/05/16/key-ai-technologies-2024-2025/ Published: 2025-05-16 Categories: technology, ai, business Tags: ai, technology, business strategy A practical guide for businesses on the key AI technologies from 2024 and emerging in 2025, detailing what they are, their importance, business utility, and limitations, all to help you make informed strategic decisions. AI isn't just that cool, experimental tech anymore. It's rocketed from "maybe someday" to "must-have now" for any serious business strategy. By the end of 2024, a whopping 49% of tech leaders said AI was already baked into their main game plan. So, let's cut through the jargon and look at the AI tech that really shook things up in 2024 and what's got us buzzing for 2025. Think of this as your cheat sheet for making smart AI moves. 2024: The Year AI Got Practical Retrieval Augmented Generation (RAG): Making AI Reliable Remember those AI "hallucinations" - when AI would just make stuff up? RAG was the 2024 superhero that swooped in to save the day. What it is: Imagine an AI that doesn't just guess based on its old training data. RAG is like giving your AI a direct line to your company's up-to-the-minute, specific information. It accesses and uses your data to give answers that are actually accurate and relevant, instead of winging it with general knowledge. Why it mattered in 2024: This became the secret sauce for trustworthy AI. By grounding AI responses in factual, company-specific data, RAG dramatically cut down on those embarrassing AI blunders. It's no wonder the RAG market hit $1.50 billion in 2024. How businesses used it: Companies started rolling out smarter customer service bots that knew the latest product specs, instantly up-to-date internal knowledge bases for staff, and AI assistants that could securely analyze proprietary company data without spilling the beans. Limitations: RAG is only as good as the data you feed it - garbage in, garbage out, as they say. Its effectiveness really leans on how good and well-organized your company's data is. Plus, setting up and maintaining these data retrieval systems can be a bit of a beast, and there are ongoing costs for keeping all that data indexed and fresh. Multimodal AI: AI That Sees, Hears, and Reads If 2023's AI was like a super-smart typist, 2024's multimodal AI is more like a perceptive colleague who can look at charts, listen to customer calls, and read reports - all at once. What it is: We're talking about AI models that aren't just text-nerds. They can process and understand a whole mix of information simultaneously - text, images, audio, and even video. Even better, they can output their responses in all of those formats too. Why it mattered in 2024: This blew the doors open for AI to work with a much richer tapestry of business data. Instead of just text, AI could get a more complete understanding of complex situations by looking at all the angles. How businesses used it: Think analyzing customer feedback from phone calls (that's your audio) right alongside their social media posts (text and images). Or creating killer marketing content that pops, and even improving product designs by actually understanding visual feedback from users. Limitations: Juggling all these different data types means things get more complicated and computationally expensive. Getting an AI to accurately line up and make sense of, say, a sarcastic comment and an accompanying eye-roll emoji, is still a tough nut to crack. Human-in-the-Loop (HITL) AI: Smarter Together Remember when everyone freaked out that AI would take all the jobs? 2024 showed us it's more about partnership than replacement. HITL became the grown-up way to use AI responsibly. What it is: These are systems where AI does the heavy lifting, crunching through tasks, but a human steps in for the really critical decisions, quality checks, or those tricky, nuanced situations AI just isn't ready for. Why it mattered in 2024: HITL was key for building trust and making sure AI was used ethically, as highlighted by industry watchers like Forbes. It's about getting AI's speed and scale, but with a vital dose of human judgment and oversight. How businesses used it: This meant safer AI in healthcare (like doctors double-checking AI's diagnostic hints), more accurate fraud detection in finance (analysts confirming AI-flagged weirdness), and better content moderation (humans making the final call on sensitive stuff). Limitations: Having humans jump in can, naturally, slow down what could be a fully automated process. It also means you've got costs for those human reviewers. For really massive volumes of tasks, scaling up the human part can be an issue, and let's not forget, humans can have biases or just get tired too. Small Language Models (SLMs): Powerful AI, Smaller Package Think of it as AI going from a room-sized mainframe to a zippy, powerful laptop. Suddenly, serious AI muscle became much more accessible. What it is: These are the leaner, meaner, more cost-effective cousins of those giant AI models. A big moment was OpenAI's launch of GPT-4o mini in July 2024, which made top-tier AI capabilities way more affordable. Why it mattered in 2024: This was huge for democratizing AI. More businesses, especially the small to medium-sized ones, could suddenly get their hands on sophisticated AI without needing a Silicon Valley budget. How businesses used it: Companies snapped these up for cost-effective AI tools to do things like summarize long documents, draft emails, or power specialized chatbots. Plus, they're easier to slot into existing software. Limitations: Smaller can mean they don't have the sheer breadth of knowledge or the raw horsepower for super-complex tasks compared to the big guys. They often still need careful fine-tuning to really nail specific business needs. Reasoning Models: AI Starts to "Think" This was when AI started to look less like a parrot and more like a problem-solver. What it is: These AI models are designed for step-by-step problem-solving and logical deduction. They're moving beyond just spotting patterns to actually figuring things out. OpenAI's o1, previewed in September and released in December 2024, was a big headline here. Why it mattered in 2024: It gave us a glimpse of AI's future potential to tackle problems that need multi-step logic - think advanced planning or deep-dive analysis. AI engineers let out a collective sigh of relief when o1-preview came out because they didn't have to spend hours telling the AI models how to reason anymore. How businesses used it (Emerging late 2024): To be honest, it was mostly hinting at future cool stuff. Think complex logistics planning, strategic financial modeling, and sophisticated scientific research. Limitations as of late 2024: This tech was (and largely still is) very new for widespread practical use. The "thinking" process chews up a lot of computing power, and there weren't many real-world business rollouts by the end of the year. 2025: AI Gets Autonomous and Integrated RAG 2.0: AI with Even Better Data Access If 2024's RAG was about getting the facts straight, 2025's RAG is about getting them faster, from more places, and with more understanding. What it is: RAG systems are getting way more dynamic. In 2025, they're plugging into real-time data feeds and even processing multimodal inputs - like understanding a product issue from an uploaded image alongside a customer's text description. The RAG market isn't slowing down, projected to hit $2.13 billion this year. Why it's important now: This deepens AI's handshake with live business information and diverse data types. The result? Richer, more current, and contextually savvy AI interactions. Business utility unfolding: Imagine AI assistants that can diagnose a customer's product problem from a photo they send, or dynamic pricing tools that react instantly to live market shifts. Or how about marketing so personalized it feels like it read your mind, based on your very latest activities. Limitations to watch: Managing all these diverse, real-time data sources adds layers of complexity. You'll need more robust infrastructure, and data privacy and security concerns get even more critical. Specialized AI Agents: Your New Digital Workforce Get ready to move beyond asking AI for information, to asking AI to do things. What they are: AI is morphing from chatbots into "AI Agents" - autonomous systems that can actually complete tasks, manage workflows, and make decisions within specific business areas. Why they're important now: This is AI graduating from an info-provider to an active "digital employee" or "colleague." It's capable of taking operational tasks and business processes off your plate. Business utility unfolding: Think automating customer service from start to finish (resolving an issue, processing a refund, and sending a follow-up email). Or proactive sales outreach triggered by specific customer behaviors, and even autonomous inventory management that just handles it. Limitations to watch: There are definite risks if these agents go rogue with too much autonomy and not enough oversight. Security is a massive concern if agents have wide access to your systems. And defining their goals, constraints, and ethical boundaries? That's a complex job. Tool Use and MCP: AI Interacting with Your Business Systems Remember USB-C making all your gadgets talk to each other? Model Context Protocol (MCP) is aiming to do something similar for AI. What it is: AI models are increasingly learning to use your existing software tools - your CRM, calendar, databases, you name it. This is happening through standardized connectors like the MCP, which Anthropic introduced in late 2024. People are already calling it the "USB-C for AI." Why it's important now: This allows AI agents to reliably get their digital hands on, retrieve, and act on real-time info from your current business systems. It makes automation more robust, scalable, and genuinely integrated. Business utility unfolding: Picture AI assistants that can book meetings directly into your team's calendars, automatically update customer records in your CRM after a call, or pull sales data from your databases to whip up a report - all seamlessly. Limitations to watch: Giving AI models API keys to your critical business systems is like handing over the castle keys - huge security implications if not managed carefully. Dealing with complex permissions and integrations across a zoo of different software can also be a headache. Advanced Reasoning Models: AI Takes on Complex Strategy The 'thinking' AIs are growing up fast. What they are: The next wave of reasoning AIs, like OpenAI's o1-pro (released March 2025), are flexing bigger muscles for complex problem-solving, multi-step logical deduction, and strategic planning. Some reasoning models are even starting to use tools and MCP while they're thinking. Why they're important now: AI is starting to get its head around more sophisticated strategic business puzzles that need deeper analytical chops and a real understanding of your business and its context. Business utility unfolding: We're looking at deeper business analytics, complex financial modeling for much better forecasting, optimizing tangled supply chains, and serious help with strategic planning and "what-if" scenario analysis. Limitations to watch: These advanced brains don't come cheap. The o1-pro, for instance, can run you up to $1 per query. They're still very computationally hungry, and getting the AI to clearly explain its complex reasoning pathway so you can actually trust it? That's still a work in progress, but we're getting there. Multi-Agent Systems: AI Teamwork on the Horizon This is where things get really futuristic - think specialized AI agents forming a committee to solve super complex problems. What they are: These are teams of AI agents, each with its own specialty, designed to collaborate on complex problems. As Botpress notes, they're for tasks too large, complex, or decentralized for a single general-purpose AI. Why they're important (for the future): This is the bleeding edge. It shows the potential for highly autonomous execution of incredibly complex tasks by distributing the AI brainpower. Business utility (Future Potential): We're dreaming here, but think potentially fully automated project management, AI-driven R&D cycles, or dynamic resource allocation across an entire company without a human lifting a finger. Some of the fast adopters in your business may already be experimenting with these ideas on a small scale. Limitations to watch: The complexity is mind-boggling. Making sure these agents coordinate, communicate effectively, and sort out their differences is a huge challenge. Debugging and keeping control over a distributed, autonomous system like this is incredibly tough. For now, it's mostly in research labs and early pilot phases - definitely not ready for prime time in most businesses. What This Means for Your Business So, what's the big takeaway? 2024 set the stage, making AI dependable and within reach. Now, 2025 is hitting the accelerator: AI is getting more independent, weaving itself deeper into our business tools, and gearing up to solve some seriously tricky problems. For anyone steering the ship, getting a grip on these changes isn't just interesting - it's how you'll spot new business opportunities, boost efficiency, and stay ahead of the pack. But let's be real: jumping into AI successfully means you've got to think hard about costs, the state of your data, keeping things secure, and getting your team skilled up. The companies that'll win aren't just those with the flashiest AI, but the ones who smartly weave these tools into how they already work and think. Nail down a clear picture of how AI can 2x progress toward your specific goals - that's your golden ticket. If you're ready to take the first steps in your AI journey, or maybe just unstuck, I've been in the trenches of AI for over 15 years and love helping organizations find their footing in this space. Drop me a line if you'd like to chat about making AI work for your business - no fancy proposals nor corporate speak required. --- ## Your AI Coding Assistant Isn't Failing. Your Management Style Is. URL: https://www.randalolson.com/2025/04/12/ai-coding-management/ Published: 2025-04-12 Categories: programming, management, ai Tags: programming, management, ai, productivity Why dismissing AI coding tools might reveal more about your management style than the technology's capabilities. Ever since I learned to program on computers, I've constantly looked for ways to optimize my workflow. From learning keyboard shortcuts to inventing my own AI automation tools, my pursuit of efficiency has bordered on obsession. But nothing has transformed my coding practice quite like AI coding assistants. These tools are revolutionizing programming, yet scroll through tech Twitter or LinkedIn and you'll find no shortage of posts declaring them useless, overhyped, or even harmful to programmers. I've noticed something interesting about these dismissive posts: they almost always come after less than a week of trying the technology. The pattern is remarkably consistent. Someone tries an AI coding tool for a week, gets frustrated when it doesn't magically solve all their problems, then writes a scathing critique that goes viral. "AI coding is all hype," they declare. "It can't even build a simple app without making basic mistakes." But here's what I've come to realize: the problem isn't with the AI tools. It's with how people are managing them. The evolving landscape of software development: AI tools are becoming essential partners in the coding process, requiring thoughtful management and direction. The Management Parallel When we hire a new programmer, we don't hand them a vague three-sentence description of a project and expect them to code the entire thing perfectly without guidance. We don't fire them when their first attempt has bugs. We don't expect them to understand our entire codebase on day one. Yet this is exactly how many people approach AI coding assistants. The truth is that effectively using AI coding tools requires the same skills as effectively managing human programmers. You need to provide clear, detailed requirements, offer necessary context about your project, review output critically, iterate on solutions, and understand the strengths and limitations of your resource. Poor managers blame their teams when projects fail due to unclear requirements. Similarly, poor AI users blame the technology when it fails to deliver on vague prompts. Common AI Coding Failures Reframed Let's look at some common complaints about AI coding tools and reframe them through the management lens: Example 1: Vague prompts → Vague requirements When someone tells an AI "create a social media app" and gets unusable code, that's equivalent to telling a junior developer "build Facebook" without any specifications. Both scenarios lead to predictable failure, but only one gets blamed on the tool rather than the communicator. Example 2: Not reviewing AI output → Not reviewing junior developer code Would you let a new hire push code directly to production without review? Of course not. Yet many people expect AI to generate perfect, production-ready code on the first try, then complain when bugs appear. Example 3: One-and-done approach → Expecting perfect first drafts Good software development is iterative. When we work with human developers, we expect multiple rounds of feedback and improvement. With AI, many users make one attempt, see flaws, and declare the entire approach worthless. Understanding these parallels is the first step on a long journey. Just like any worthwhile skill, learning to manage AI coding assistants has a learning curve. Learning Curve Reality I started heavily using AI coding tools in 2023. Like many, my first experiences were basic: using ChatGPT to write small code snippets or modify existing code. My initial attempts with GitHub Copilot were frustrating and unproductive. The turning point came when I started using Cursor AI. I had an "aha" moment when I discovered it could analyze an entire codebase that I was onboarding onto, create architecture diagrams, and explain the code base to me in plain English after a few minutes. This discovery transformed how I approached new projects and legacy codebases. Cursor AI in action: demonstrating its ability to analyze and explain complex codebases. Since then, I've been learning new techniques daily, building an ever-growing knowledge base for effectively managing AI coding agents. Each day brings new discoveries about how to better direct these tools. Each day involves practice and patience, just the same as if I were managing people. This learning process takes time, just like learning any other programming skill. We didn't expect to master C++ in a week, so why do we expect to master AI-assisted programming in the same timeframe? The Meta-Programming Mindset The rise of AI coding tools represents a fundamental shift in how we approach software development. We're moving from programming to meta-programming: directing AI systems that can handle much of the implementation work rather than writing every line ourselves. This shift demands a different set of skills. The most successful AI coders approach these tools as technical managers and collaborators rather than expecting magical code generation. They understand that their role is to provide direction, context, and critical evaluation, not to passively consume whatever the AI produces. This explains why those viral dismissals of AI coding tools are so revealing. When someone declares AI coding "useless" after a brief trial, they're inadvertently demonstrating their own management limitations. They haven't recognized that effectively using these tools requires the same skills as effectively leading a development team. The developers who will thrive in the AI era are those who embrace this meta-programming mindset and invest in developing their AI management skills. The rewards—dramatically increased productivity, better code quality, and the ability to tackle more ambitious projects—are well worth the learning curve. Tips for AI Coding Management Drawing from both personal experience, here are some crucial tips for effectively managing your AI coding assistant: 1. Context management is vital Just as you wouldn't expect a new team member to understand your codebase without proper onboarding, your AI coding assistant needs proper context to be effective. Before asking it to write or modify code, explicitly tell it which files to examine. If you're unsure about the exact files, instruct it to search for relevant code, functions, or patterns first. This approach prevents the AI from writing redundant code or making changes that conflict with existing implementations. 2. Stay current with web search AI models have training cutoff dates, which means they might not know about the latest changes to the libraries you're using or best practices. Modern AI coding environments like Cursor provide web search capabilities (via `@web`) and can reference specific documentation (like `@openai` for OpenAI documentation). Use these features to ensure your AI assistant is working with the most up-to-date information. This is particularly crucial when working with rapidly evolving frameworks and libraries. 3. Treat AI-assisted coding as a conversation Resist the urge to immediately dive into coding. Instead, treat your AI coding assistant as a collaborative partner. Start by discussing options, gathering its thoughts on your feature requirements or bug reports, and developing a solid plan together. This conversational approach often reveals potential issues or alternative solutions you might not have considered. Once you've established a clear direction, then proceed with implementation. These practices transform your interaction with AI coding tools from a simple code generation exercise into a truly collaborative development process. The key is to remember that you're not just using a tool; you're managing an intelligent assistant that becomes more effective the better you communicate with it. Want to Become a Better AI Manager? I teach workshops on effective AI coding techniques, helping teams develop the skills needed to maximize their productivity with these powerful tools. If your organization wants to stay ahead of the curve and ensure your developers are getting the most out of AI coding assistants, let's talk. After all, the difference between a frustrating AI experience and a revolutionary one often comes down to one thing: management. --- ## How to beat the Original I.Q. Tester URL: https://www.randalolson.com/2024/03/26/how-to-beat-original-iq-tester/ Published: 2024-03-27 Categories: puzzles, strategy, analysis Tags: puzzles, strategy, analysis Randy Olson shows you how to beat the Original I.Q. Tester during a trip down memory lane. Growing up, there was a simple yet mesmerizing game that captured my attention every weekend known as the Original I.Q. Tester. This game, often found at the tables of family restaurants among syrup bottles and sticky menus, consisted of a wooden triangle filled with pegs, challenging players to jump pegs over each other until only one remained. The game became a ritualistic endeavor for me, especially on Sundays when I visited IHOP with my mom and step-dad on our weekly family outing. Amidst the aroma of pancakes and coffee, this game was one of my early forays into problem-solving and strategy, igniting a curiosity that has stayed with me into my later life and career. Maybe you have fond memories of playing this game, too. The nostalgic Original I.Q. Tester While cleaning out my mom's house earlier this month, an unexpected treasure surfaced from the depths of a moving box: an old Original I.Q. Tester. Had she kept it as a keepsake of our Sundays together, or did she intend to surprise me with it one day and then forgot? Either way, this rediscovery reignited my resolve to master the puzzle, not just through trial and error, but this time by harnessing the power of computation. For what other reason did I get this PhD, after all? How difficult is Original I.Q. Tester? To tackle the puzzle's challenge, I created a simulator for the puzzle in Python, a tool that modeled every conceivable starting position — each variation of the initial empty peg — and traced every potential path the game could follow. This digital experiment unearthed a staggering 7,335,390 possible outcomes, which highlights the game's deceptively complex nature. To put that number into context, even if I was able to play 10 unique paths every Sunday, it would have taken me over 14,000 years' worth of Sundays to play every possible path. That's a lot of pancakes! Interestingly, the data from my simulator revealed that about 94% of all possible games end with two or more pegs defiantly remaining on the board, a statistic that underscores the puzzle's difficulty and its knack for humbling even the most strategic minds — mine among them. Where's the best starting point? In my initial encounters with the I.Q. Tester, I instinctively chose to start with an empty peg somewhere in the center, often in the middle of the third row from the top, assuming this offered the greatest flexibility for interesting maneuvers. This intuition seemed logical, suggesting a wealth of strategic paths from the heart of the board. However, after rigorously analyzing every conceivable game trajectory through my simulator, I was amused — and admittedly a bit chagrined — to find that starting in the middle actually set the stage for the worst possible outcomes, dramatically limiting the paths that could lead to a single peg standing. The data compellingly pointed to the edges of the third row from the top and the center of the bottom row as the optimal starting points, revealing a counterintuitive strategy (for me) that significantly increases your chances of solving the puzzle. What's a good jump strategy to reach 1 peg? Next, I wanted to learn what sort of jump strategy worked best for the puzzle. In my early days, I oscillated between instinctive play, selecting jumps on a whim, and a more deliberate strategy, aiming to consolidate my pegs towards the board's center in hope of increasing the availability of subsequent jumps. This latter approach was grounded in a belief that a concentrated cluster of pegs would inherently offer more opportunities for strategic moves. The data from my game-path analysis painted a clear picture: consistently opting for moves that draw pegs closer significantly boosts the likelihood of achieving the coveted outcome of a solitary peg standing. Conversely, the allure of seemingly strategic, yet ultimately divisive jumps proved to be a pitfall, as spreading the pegs across the board invariably diminishes your chances of success. Notably, early on it's fairly difficult to put this strategy into effect because the pegs are everywhere. That means you have to focus on eliminating one side of the triangle first so you can later concentrate the pegs on one side of the puzzle. This revelation was a nod to my initial instincts, affirming that cohesion, not dispersion, was key to conquering the game's intricate puzzle. Just like so many things in life, right? And sure enough, when I followed this strategy myself through a few games, I was finally able to end a game with a single peg. Victory at last! Some example games If you're curious what this strategy looks like when played out, here's one of the 439,000 game paths that lead to a single peg. Click to see a 1-peg ending Board labels: A1 B1 B2 C1 C2 C3 D1 D2 D3 D4 E1 E2 E3 E4 E5 Initial board state (1 = peg, 0 = empty): 1 1 1 0 1 1 1 1 1 1 1 1 1 1 1 Move 1: A1 to C1 0 0 1 1 1 1 1 1 1 1 1 1 1 1 1 Move 2: D1 to B1 0 1 1 0 1 1 0 1 1 1 1 1 1 1 1 Move 3: C3 to C1 0 1 1 1 0 0 0 1 1 1 1 1 1 1 1 Move 4: E5 to C3 0 1 1 1 0 1 0 1 1 0 1 1 1 1 0 Move 5: B1 to D1 0 0 1 0 0 1 1 1 1 0 1 1 1 1 0 Move 6: B2 to D4 0 0 0 0 0 0 1 1 1 1 1 1 1 1 0 Move 7: E1 to C1 0 0 0 1 0 0 0 1 1 1 0 1 1 1 0 Move 8: E3 to E5 0 0 0 1 0 0 0 1 1 1 0 1 0 0 1 Move 9: C1 to E3 0 0 0 0 0 0 0 0 1 1 0 1 1 0 1 Move 10: E2 to E4 0 0 0 0 0 0 0 0 1 1 0 0 0 1 1 Move 11: E5 to E3 0 0 0 0 0 0 0 0 1 1 0 0 1 0 0 Move 12: E3 to C3 0 0 0 0 0 1 0 0 0 1 0 0 0 0 0 Move 13: C3 to E5 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 Similarly, if you want to see what the opposite of this strategy looks like, here's a typical game path that ends with 3 pegs on the board. Click to see a 3-peg ending Board labels: A1 B1 B2 C1 C2 C3 D1 D2 D3 D4 E1 E2 E3 E4 E5 Initial board state (1 = peg, 0 = empty): 1 1 1 0 1 1 1 1 1 1 1 1 1 1 1 Move 1: C3 to C1 1 1 1 1 0 0 1 1 1 1 1 1 1 1 1 Move 2: E3 to C3 1 1 1 1 0 1 1 1 0 1 1 1 0 1 1 Move 3: D1 to D3 1 1 1 1 0 1 0 0 1 1 1 1 0 1 1 Move 4: D4 to D2 1 1 1 1 0 1 0 1 0 0 1 1 0 1 1 Move 5: B2 to D4 1 1 0 1 0 0 0 1 0 1 1 1 0 1 1 Move 6: E2 to C2 1 1 0 1 1 0 0 0 0 1 1 0 0 1 1 Move 7: B1 to D3 1 0 0 1 0 0 0 0 1 1 1 0 0 1 1 Move 8: D4 to D2 1 0 0 1 0 0 0 1 0 0 1 0 0 1 1 Move 9: C1 to E3 1 0 0 0 0 0 0 0 0 0 1 0 1 1 1 Move 10: E4 to E2 1 0 0 0 0 0 0 0 0 0 1 1 0 0 1 Move 11: E1 to E3 1 0 0 0 0 0 0 0 0 0 0 0 1 0 1 Finally, if you just want to feel better about yourself for never passing the 3-peg mark, here's one of the 6 ways to fail so spectacularly that you end up with 10 pegs left on the board and no valid jumps. Click to see a 10-peg ending Board labels: A1 B1 B2 C1 C2 C3 D1 D2 D3 D4 E1 E2 E3 E4 E5 Initial board state (1 = peg, 0 = empty): 1 1 1 1 0 1 1 1 1 1 1 1 1 1 1 Move 1: E4 to C2 1 1 1 1 1 1 1 1 0 1 1 1 1 0 1 Move 2: B1 to D3 1 0 1 1 0 1 1 1 1 1 1 1 1 0 1 Move 3: E2 to C2 1 0 1 1 1 1 1 0 1 1 1 0 1 0 1 Move 4: D3 to B1 1 1 1 1 0 1 1 0 0 1 1 0 1 0 1 Take-aways Embarking on this computational quest to solve the Original I.Q. Tester was not just an exercise in tryhardism or analytics; it was a bridge to the past for me, a connection to the cherished moments I used to spend with my mom. I found myself often thinking back to those Sunday mornings, her laughter and gentle encouragement mingling with the clatter of pegs while we talked about our upcoming week and I invariably moaned about something only a teenager would care about. I miss her deeply, yet through this project, I've found solace in the realization that our loved ones often leave us with puzzles — not just those of wood and pegs, but puzzles of the heart, teaching us about resilience, togetherness, and the beauty of simple shared moments. I hope this story inspires you to cherish the puzzles in your life, recognizing them as opportunities for growth, reflection, and a reminder of the bonds that, though unseen, forever hold us close to those we love. Or, at the very least, you can impress someone with your newfound puzzle-solving skills. If you'd like to chat puzzle strategy or even share your own memories, feel free to contact me anytime. --- ## Things I can lift: How I visualize my strength training progress URL: https://www.randalolson.com/2021/04/15/things-i-can-lift-week7/ Published: 2021-04-15 Categories: data visualization, personal Tags: data visualization, humor, strength training Randy Olson demonstrates a humorous yet useful way to visualize strength training progress. Earlier this year, I set up a power lifting cage for myself and restarted my journey of rapidly building strength. Every other day I perform squats, once a week I perform a deadlift, I alternate between bench presses and overhead presses, and I throw in pull-ups and chin-ups on the non-deadlift days. It's a fairly standard Starting Strength program. The difference this time is that I decided to track and visualize my progress every day in a spreadsheet. Part of the fun of rapid strength training is keeping track of your progress, and I like to take that to the next level by keeping track of things that I'm newly capable of lifting. It's a lot more fun to say you can squat a refrigerator than to say you squat 200 pounds, right? My week 7 progress on 5-rep lifts One of my early goals has been to get to a point where I can replicate the Dirty Dancing scene in Crazy, Stupid, Love: I'll never share Ryan Gosling's charm, but at least I'll share his strength Some other exciting milestones for me include being able to bench press a full beer keg and lift a medium-sized refrigerator with relative ease. I would've thought that refrigerators weighed way more than ~200 pounds! The ultimate milestone, of course, has been reaching the point where I can lift the legend himself: Arnold Schwarzenegger. At his prime bodybuilder weight, Arnold weighed in at about 225 pounds, which is ludicrous to imagine because of how muscular he is. Henceforth I'm going to measure multiples of 225 pounds as "Arnolds." This is all tongue-in-cheek, of course, and not meant to be taken too seriously. I'll make sure to provide an update down the line when I've hit some more big milestones. As the legend says: What object-lifting milestones should I aim for next? --- ## A data-driven look at marble racing URL: https://www.randalolson.com/2020/05/24/a-data-driven-look-at-marble-racing/ Published: 2020-05-24 Categories: analysis, data visualization Tags: data analysis, data visualization, marble racing Randy Olson looks at the new marble racing phenomenon through the eyes of a data scientist. Do marbles have measurable "skill"? Race tracks. High stakes. Rabid fans. Marbles...? Marble racing has taken the world by storm in 2020. Unless you've been living under a rock -- which is entirely understandable given the state of the world -- you've probably seen marble racing in one form or another in the past few months. If not, you owe it to yourself to watch one of the most satisfying comeback stories of the year: Race pic.twitter.com/POe5ujIQ5a — viral posts (@BestViralPosts) January 29, 2020 Jelle's Marble Runs started as a quirky YouTube channel back in 2006 and has refined the art of marble racing to the point that many --- including sponsor John Oliver from Last Week Tonight --- consider marble racing a legitimate contender for the national sports spotlight. Given that Jelle's Marble Runs just completed their popular Marbula One competition last month, I was curious to look at the race results to see if these races were anything more than chaos. Do some marbles race better than others? Who would I put my money on in season 2 of Marbula One? Is one of the marbles having an affair? If any of these questions interest you, read on and I'll answer some of them. The first step to answering these questions was to get some data. Thankfully, all of the Marbula One videos are organized in a YouTube playlist available here. From every race, my marble racing analytics team recorded each marble racer's qualifier performance, total race time, average lap time, final rank, and some other statistics. That dataset is available for download on my website here. Do some marbles race better than others? At first thought, it might seem like a ridiculous question to ask whether one marble racer is more skilled than another. All marbles are created equal, after all, and there's not much to differentiate one marble from another. However, it's clear that not all marbles trained equally in preparation for the first season of Marbula One. The chart below shows the distribution and median of race times of every individual marble racer, sorted from fastest to slowest race times. Note that I had to standardize their race times because each of the race tracks took a varying amount of time. Thus, a marble's performance in one race is measured as how much faster or slower they were than the average time it took all of the marbles to complete that race track. Starry clocked in the greatest individual race time this Marbula One season, crossing the finish line a full 7 seconds faster than the bulk of the pack in the first race of the season. Unfortunately for Team Galactic, Starry appeared to be fatigued after this race as she turned in mediocre performances for the rest of the season. Outlier performances aside, we see something that's rather surprising in this data: Several marbles seemed to consistently perform above average in season 1 of Marbula One. Snowy, Smoggy, Speedy, and Prim stand out as the top racers, consistently vying for the top spot in every race. Speedy especially demonstrated why he's considered one of the sport's top athletes with stellar performances in every Marbula One race. On the other end of the spectrum --- and equally surprising --- we see that some marbles consistently perform worse than their competitors. Mary, Sublime, Vespa, Snowflake somehow always found themselves behind everyone else, with Mary even failing to finish one race because she fell a full lap behind. Let's look into how these performances affected their teams. What marble teams should I bet on? If you're new to marble racing and looking for a took to root for, I've compiled a chart with you in mind. Below I show the distribution of each team's placement in every Marbula One race, sorted from the best-ranking teams to the worst. If you like to root for the winners, pick a team near the top of this chart. If you like to root for the underdogs, pick a team near the bottom of this chart. The Savage Speeders, one of the most decorated teams in Marble League history, lived up to their name this season and consistently ranked in the top 8. The lone exception for the Savage Speeders was one out-of-character performance in race 5 from Rapidly, where some suspect he put in a late night at the clubs the night before the race. The Savage Speeders will no doubt ride on this victory into the 2020 Marble League this summer, as they expect to easily pass the qualifiers and add to their mountain of gold medals. The Hornets, on the other hand, are performing about as well as the NBA team they share a name with. As a relatively new team to the Marble League, the Hornets are still trying to make a name for themselves as they hope to avoid relegation. Luckily for them, the Hornets were officially invited to participate in this year's Marble League and will have the opportunity to prove that they belong in the big leagues. Interestingly, teams such as the Snowballs and Team Primary had an up-and-down Marbula One season, sometimes placing first in one race then placing last in the next --- much to the disappointment of their fans. Let's see why these teams seemed to struggle. Some marble racing teams need a reshuffle To get a better idea of why some teams were so hot and cold during throughout the Marbula One season, I matched each team up on the line plot below. Each team had two athletes competing in this season, and I compared the average race rank of each team's athletes. As an example for team Snowballs, Snowy ranked 3rd in every Marbula One race on average vs. Snowflake's average rank of 13. From the above chart, it's clear that some teammates are holding their team back. Team Snowballs and Team Primary could have been in contention for the podium if it weren't for Snowflake and Mary, and Sublime embarrassed his team in every race he participated in --- including a race he didn't even finish. What's worse is that Snowflake, Mary, and Sublime are all team captains, who you expect to lift their team up instead of hold them back. It might be time for Snowflake and Mary to consider stepping back into managerial roles. It's been years since either of them have stood on the podium for individual events, and it's hard to deny that their recent performances are holding their teams back. Does qualifier performance matter? The last question I wanted to investigate was brought to me by a fellow marble racing fan: Does qualifier performance matter? In Marbula One races, before the main race every marble completes one lap around the race track. Their performance in that one lap determines their starting position in the upcoming race, where the marble racer with the fastest qualifier lap starts ahead of everyone else and the marble racer with the slowest qualifier lap starts behind everyone else. Is a marble racer's fate sealed if they don't perform well in the qualifier? To answer this question, I matched up each marble racer's qualifier rank (x-axis) with their final rank in the corresponding race (y-axis). To make the trend a little clearer, I binned the ranks into the Top 4, Ranks 5 - 8, Ranks 9 - 12, and the Bottom 4. In general, if a marble racer places in the top 50% in the qualifier, she will place in the top 50% in the race as well. The opposite case is true as well when a marble racer places in the bottom 50% of the qualifier. It's safe to say that qualifiers generally decide who will bring home points for their team in the upcoming race. However, it's not unheard of for marble racers to defy the odds. Snowy amazingly claimed the gold in race 6 despite a poor showing in the qualifier, and Wospy infamously failed to even complete race 2 after an impressive showing in the qualifier. No matter the cause, fans live for these kinds of upsets! The near future of marble racing In this post, I was quite surprised to discover that not all marble athletes are created equal, and that some marbles seem to consistently perform better or worse than their peers. It's possible that this is all statistical noise, but we'll have to wait for Marbula One season 2 sometime in Autumn 2020 to find out. Given the track record of some of these marble athletes, I suspect they will once again find themselves ahead of the pack in season 2. If you're excited for more marble racing, thankfully you won't have to wait very long. This year's Marble League starting in late June: and there are plenty more events to catch up with on their YouTube here. Enjoy, marble racers! --- ## Does batting order matter in Major League Baseball? A simulation approach URL: https://www.randalolson.com/2018/07/04/does-batting-order-matter-in-major-league-baseball-a-simulation-approach/ Published: 2018-07-04 Categories: analysis, python, statistics Tags: analysis, baseball, data visualization, major league baseball, python Randy Olson uses data science to learn whether batting order matters in Major League Baseball. If you've ever watched Major League Baseball, one of the feature points of the sport is the batting line-up that each team decides upon before each game. Traditional baseball logic tells us that speedy, reliable hitters like Trea Turner should lead the line-up of batters; slower, power hitting juggernauts like Giancarlo Stanton should make up the middle of the line-up; and the less-than-stellar, oh-god-why-is-this-person-even-trying sluggers like Bartolo Colon should fill in the back of the order. Perhaps you've wondered, like myself: How did we arrive at these rules of thumb? Since every batter gets several chances at the plate anyway---and it's not like we start every inning at the top of the batting order---does batting order even matter? Many pitchers, like Bartolo Colon, aren't known to be very good at batting. In today's post, I want to take a simulation-based approach to learn whether batting order matters in Major League Baseball. If we hold everything equal, and all of the batters in our batting line-up hit equally well, will our team perform better if we design the batting order in a particular way? If we introduce a star batter to that team, does it matter where he lines up? Or if we introduce a poor batter to that team, does their position in the line-up matter? In the following sections, I'm going to walk through several data visualizations that summarize the results of 55 million simulated baseball games. These are all simplified baseball simulations, of course, as I don't have the funding nor development effort of a game studio to replicate a full baseball simulation. Regardless, I've simulated baseball games where 9 batters make successful hits according to a pre-determined Batting Average (BA) and make singles (64% of all hits), doubles (20%), triples (2%), and home runs (14%) according to league averages. When 3 outs are made in an inning by the defense, the bases are cleared and the next inning picks up in the batting order where the previous inning left off. All the while during each 9-inning simulation, I track batting statistics and total runs scored by the team, which I visualize below. If you'd like to take a look at the simulation underlying these visualizations---or better yet, if you'd like to improve the simulation by contributing some code---you can find the Python simulation code on my GitHub repository here. Does it matter where the good batter lines up? If we have 8 average batters and 1 exceptionally good batter, does it matter where the good batter lines up? Say we have Nelson Cruz lining up with a team of Daniel Descalso clones, where in the batting order would you place Cruz? In this simulation, I created a line-up of 8 clone batters and 1 exceptionally good batter that stands in as the "Designated Hitter" (DH) hitting a 0.35 BA. The 8 clone batters are assigned a range of batting performance on the y-axis, ranging from extremely poor (0.1 BA) to exceptionally good hitting (0.35 BA). In the following chart, we look at the relative number of runs scored by the team according to the corresponding team BA and the DH's batting position. On the x-axis, we see the effect of moving the DH to different positions in the line-up. 1.0 (white) means that the team achieves an average number of runs, below 1.0 (orange) means that the team achieves fewer runs, and above 1.0 (purple) means that the team achieves more runs. Note that these values are normalized separately for every Team BA, so we should only compare between DH batting positions here, and not between Team BAs. As we might expect, teams with a higher BA will always score more runs, so those comparisons are pointless. The effect of batting order is immediately noticeable: The better the DH hits compared to their team, the more important it is for the DH to line up near the beginning of the batting order. Interestingly, when the team bats nearly as well as the DH (Team BA=0.3 and higher), batting order doesn't matter at all. (For the stats nerds: None of the differences across BA=0.35 are statistically significant according to a Wilcoxon rank sum test.) Yet as the disparity in batting ability increases, teams that place the DH at the top of the batting order will score more runs over the season. Does it matter where the pitcher lines up? What if we have the opposite situation where we have an exceptionally poor batter? Say we have Bartolo Colon lining up with that same team of Daniel Descalso clones, where in the batting order would you place Colon? (My favorite answer so far: Make 9 Bartolo Colon clones and watch the hilarity that ensues when they try to make a hit.) In this simulation, I created a line-up of 8 clone batters and 1 extremely poor batter that stands in as the "Pitcher" hitting a 0.1 BA. The 8 clone batters are assigned a range of batting performance on the y-axis, ranging from extremely poor (0.1 BA) to exceptionally good hitting (0.35 BA). In the following chart, we look at the relative number of runs scored by the team according to the corresponding team BA and the Pitcher's batting position. On the x-axis, we see the effect of moving the Pitcher to different positions in the line-up. Again, 1.0 (white) means that the team achieves an average number of runs, below 1.0 (orange) means that the team achieves fewer runs, and above 1.0 (purple) means that the team achieves more runs. Note that these values are normalized separately for every Team BA, so we should only compare between Pitcher batting positions here, and not between Team BAs. As we might expect, teams with a higher BA will always score more runs, so those comparisons are pointless. Now we see the complete opposite effect of the DH: The worse the Pitcher hits compared to their team, the more important it is for them to hit last in the batting order. To summarize these two discoveries, I ran one more set of baseball simulations where 8 batters hit with a BA of 0.25, and a "Hitter" has a varying BA (y-axis) and batting position (x-axis). You can read the chart below the same as the previous two charts. It seems the "lead with your best, finish with your worst" baseball mantra is supported by these simulations. However, it's important to note the magnitude of the effect here. In the case of the worst Hitter (BA=0.1), the team that places the poor Hitter last in the batting order will score about 2.7% more runs in a season than the team that leads with the poor Hitter. Likewise in the case of the best Hitter (BA=0.35), the team that leads with the best Hitter will score about 1.6% more runs in a season than the team that finishes with the best Hitter. In both cases, the increased runs scored are statistically significant (according to a Wilcoxon rank sum test, p<1e-5), but the effects of batting order are fairly small. Regardless, in any competitive sport you need to take advantage of any benefit that you can gain over your competition---and a dozen or so extra runs can mean a few more wins that year---which is why so much effort goes into designing batting orders in Major League Baseball. Why does batting order matter? In the above sections, hopefully I've convinced you that batting order matters when you have an exceptionally good or poor batter on your team. Yet perhaps the biggest question of all remains: Why does batting order matter? Every batter has the same opportunity to make their contributions at the plate, and in the case of my simulations, everyone except 1 batter hits with exactly the same BA. What's going on? To answer that question, we return to the first set of simulations where I simulated a team of equally-performing batters. This time, we're working with a full line-up of Daniel Descalso clones: All reliable contributors hitting a 0.25 BA, but none of them standing out on the team. If we kept all of the clones at the same position for 1 million games and batting order didn't matter, no clone should perform better than the other clones, right? Wrong. The simple answer is that not every batter in the line-up has the same opportunity to contribute, even in the case where all 9 batters hit with exactly the same BA. Barring special circumstances, all batters are more-or-less guaranteed 3 At Bats per game, even in the case where no batter makes a successful hit. However, as soon as one of the batters makes a successful hit, the first batter in the line-up is guaranteed a 4th At Bat. Same thing for the second batter in the line-up when the next hit is made, and so on. These small differences in At Bat opportunities add up to the point that Daniel Descalso clone #1 will have roughly 142 more At Bats than Daniel Descalso clone #9 over the course of the season---nearly 1 additional At Bat per game in a season! The findings in the earlier sections should all make sense now: If we have an exceptionally good batter, we want them to bat as much as possible. Likewise, if we have an exceptionally poor batter, we want them to bat as little as possible. What about middle batters? If At Bat opportunities were the whole answer to the batting order question, then what about middle batters? If all we care about is giving our best batters the most At Bat opportunities, why bother with placing heavy home run hitters at the 4th and 5th positions? Of course, At Bat opportunities aren't the whole answer; it also matters how many teammates are on base when the batter comes up to the plate. If we analyze the same 1 million games from the above section, we'll make an interesting finding: Daniel Descalso clone #4 will see, on average, more of his teammates on base when he goes up to bat than any other Daniel Descalso clone. As a direct consequence, Daniel Descalso clone #4 will also contribute more RBI per game: This advantage to Daniel Descalso clone #4 is primarily conferred from the 1st inning, where he's highly likely to have a teammate on base if he goes to bat in the 1st inning. Daniel Descalso clone #5 and beyond will have a similar guarantee, but it becomes increasingly unlikely to make it beyond the 4th batter in the 1st inning. After the 1st inning, there are no such guarantees in the batting order. Thus, it benefits the team to put a batter in the 4th position that is most likely to score a RBI in the 1st inning---even if their teammate is on 1st base. In other words: a power hitter that can hit home runs. Although the effect may again seem small, Daniel Descalso clone #4 will contribute 8 more RBI than Daniel Descalso clone #9 over the course of the season, even in these simulations where all 9 clones are batting with the exact same BA. As such, the 4th position is another area that can be planned around to maximize a team's run-producing potential, and any competitive team should take advantage of that fact. There's an additional layer of strategy beyond what I've simulated here. In the real world, not all players bat equally well, and teams will fill the first 3 batting positions with high BA batters that can get on base. This strategy only amplifies the RBI-producing potential of the 4th batter, and tends to produce more chains of successful hits that set up the 4th batter with players on base. For now, I'll leave those simulations for a future article. Bonus: Grand Slams While analyzing the 1 million games from the previous section, I also wanted to take a look at Grand Slam probability. For a batter to hit a Grand Slam, their teammates must fill the bases and leave at least 1 out remaining, and the batter must hit a home run. It should be no surprise that Grand Slams are incredibly rare in baseball, and make for exciting and game-altering plays. For example, David Ortiz's Grand Slam in the 8th inning of game 2 of the 2013 ALCS brought the Red Sox back into the game from a 5-1 deficit, and made it possible for the Red Sox to clutch victory from the jaws of nearly guaranteed defeat. Interestingly, if we take a look at the batting statistics from the Daniel Descalso clone simulations, we find that Daniel Descalso clone #6 will walk up to more Bases Loaded situations than his teammates: And following from that, Daniel Descalso clone #6 will hit more Grand Slams than his teammates: Over time, Daniel Descalso clone #6 is roughly 40% more likely to hit a Grand Slam than his teammates that hit with the exact same BA. Grand Slams are so rare that it's likely not worthwhile to design a batting order around this finding, but I found it surprising that even Grand Slam probability is affected by batting order. So there you have it: Batting order definitely matters in Major League Baseball, and even basic simulations confirm conventional baseball wisdom. There's a lot more that we can learn from these baseball simulations, but I'd like to hear your thoughts first. I hope you enjoyed reading this article as much as I enjoyed working on it. If you'd like to run your own baseball simulations, you can find my Python code on GitHub here. If you have any comments or questions about the simulations, add them to the comments below. --- ## Traveling salesman portrait in Python URL: https://www.randalolson.com/2018/04/11/traveling-salesman-portrait-in-python/ Published: 2018-04-11 Categories: data visualization, python, tutorial Tags: data visualization, optimization, python, traveling salesman problem, tutorial Randy Olson shows how you can create your own traveling salesman portrait using Python. Last week, Antonio S. Chinchón made an interesting post showing how to create a traveling salesman portrait in R. Essentially, the idea is to sample a bunch of dark pixels in an image, solve the well-known traveling salesman problem for those pixels, then draw the optimized route between the pixels to create a unique portrait from the image. Antonio is a fan of Frankenstein, so he created a traveling salesman portrait from an old Frankenstein image. I liked the idea of the traveling salesman portrait, so I thought it would be a fun exercise to re-create it in Python. Below, I walk through the code line-by-line. If you want the full code snippet, you can find it on my personal projects GitHub repository. To start, we need an image of someone. For ease of comparison, I decided to use the same image as Antonio. Franky is looking handsome as ever. import urllib.request import os image_url = 'http://www.randalolson.com/assets/2018/04/Frankenstein.jpg' image_path = 'Frankenstein.jpg' if not os.path.exists(image_path): urllib.request.urlretrieve(image_url, image_path) Note: We can use any image we want, but this algorithm works best for images with light backgrounds. Next, we need to convert that image to black and white. PIL makes this operation pretty straightforward. from PIL import Image original_image = Image.open(image_path) bw_image = original_image.convert('1', dither=Image.NONE) bw_image Now we can use NumPy to identify the black pixels and select a random subset of them: import numpy as np import matplotlib.pyplot as plt bw_image_array = np.array(bw_image, dtype=np.int) black_indices = np.argwhere(bw_image_array == 0) # Changing "size" to a larger value makes this algorithm take longer, # but provides more granularity to the portrait chosen_black_indices = black_indices[ np.random.choice(black_indices.shape[0], replace=False, size=10000)] plt.figure(figsize=(6, 8), dpi=100) plt.scatter([x[1] for x in chosen_black_indices], [x[0] for x in chosen_black_indices], color='black', s=1) plt.gca().invert_yaxis() plt.xticks([]) plt.yticks([]) Now all that's left to do is solve TSP for those 10,000 pixels. To do that, we first have to define the distance between every pixel. In this case, we're going to define distance between two pixels as the Euclidean distance between their x,y coordinates in the image. With that definition in mind, we can calculate the distances between all 10,000 pixels: from scipy.spatial.distance import pdist, squareform distances = pdist(chosen_black_indices) distance_matrix = squareform(distances) Great! The result is a giant 10,000 x 10,000 matrix with the Euclidean distances between every pixel. Now we can provide that matrix to an off-the-shelf traveling salesman problem solver: from tsp_solver.greedy_numpy import solve_tsp optimized_path = solve_tsp(distance_matrix) optimized_path_points = [chosen_black_indices[x] for x in optimized_path] plt.figure(figsize=(8, 10), dpi=100) plt.plot([x[1] for x in optimized_path_points], [x[0] for x in optimized_path_points], color='black', lw=1) plt.xlim(0, 600) plt.ylim(0, 800) plt.gca().invert_yaxis() plt.xticks([]) plt.yticks([]) And voilà! We now have a traveling salesman portrait of the ever-handsome Frankenstein. Finally, a side-by-side comparison: plt.figure(figsize=(16, 10), dpi=100) plt.subplot(1, 2, 1) plt.imshow(original_image) plt.grid(False) plt.xlim(0, 600) plt.ylim(0, 800) plt.gca().invert_yaxis() plt.xticks([]) plt.yticks([]) plt.subplot(1, 2, 2) plt.plot([x[1] for x in optimized_path_points], [x[0] for x in optimized_path_points], color='black', lw=1) plt.grid(False) plt.xlim(0, 600) plt.ylim(0, 800) plt.gca().invert_yaxis() plt.xticks([]) plt.yticks([]) If you want to make your own traveling salesman portrait, you can use my Python script on GitHub. Enjoy! --- ## How many college football teams can you watch in-person in one football season? URL: https://www.randalolson.com/2018/03/20/how-many-college-football-teams-can-you-watch-in-person-in-one-football-season/ Published: 2018-03-20 Categories: machine learning Tags: artificial intelligence, evolutionary computation, genetic algorithm, machine learning, optimization, road trip, traveling salesman problem, united states, usa Randy Olson uses machine learning to discover how many college football teams can you watch in-person in one football season, which results in a 3 1/2 month long trip around the USA. The college football season once again came to an end back in January, which now leaves us college football fans with nothing to do but post football memes online and feign interest in other sports until August rolls around again. This year, I've decided to jump the gun and start planning ahead for the 2018 college football season. With the 2018 college football season schedule already announced, we can start planning our vacations to see our favorite teams, rivalry games, and stadiums. But what if we did something a little different this year? What if we embarked on a trip that was a little more... epic in scope? Ohio State's stadium, nicknamed "The Horseshoe," is a must-visit for any college football fan If you've followed my blog over the years, you'll know that I'm no stranger to planning epic, once-in-a-lifetime road trips. This time, I'm going to make a trip that allows us to watch as many different college football teams in-person as possible during the 2018 regular season. If you don't care about how I made this trip, skip down to the "Optimized college football trip v1.0" section. Rules of the trip The best place to start when planning a trip of this magnitude is to set a few ground rules. In this case, I'm going to make a handful of assumptions: We only care about Division I FBS teams for this trip. 130 teams are enough to work with already. We especially care about Power 5 conference teams. Most of the best college football teams are in the Power 5 conferences, so when push comes to shove, we prefer seeing a Power 5 team over any other. We want to see as many big rivalry games as possible. Big rivalry games are unique and fun to watch. We will attend a maximum of one game per day. Without advance knowledge of what time the games are taking place, it's impossible to schedule multiple games per day. We will need to fly between every game. Sometimes driving between the games may be possible---especially between Saturday and Thursday games---but for the most part we will need to hop on a plane to make the next game on time. We these rules in mind, we can now formulate a plan on how to create this trip. How to optimize the trip If you're familiar with college football, optimizing this trip is no small feat: As of the 2018 season, there will be 130 Division I football teams with 51 days of games scheduled over a 3 1/2 month period, typically on Thursdays, Fridays, and Saturdays. Each of the 51 days have somewhere between 1 and 77 games scheduled on that day, and it's our task to choose a series of 51 games that allows us to see as many teams as possible. The challenge comes when we calculate the number of possible trips we can choose: On August 25, there are only 4 games to choose from. Easy enough. On August 30, there are 9 different games to choose from, for a total of (4x9=) 36 possible trips. On August 31, we add another 5 games to choose from and we're up to (4x9x5=) 180 possible trips. By day 51, we're up to roughly 1.6 x 1032 possible trips. That got out of hand quickly, didn't it? Even with a computer, solving this problem in a brute force manner isn't practical. On my computer, I can evaluate about 1,000 different trips per second. Even if I spread this program across 1,000,000 more computers, it would take about 5 x 1015 years to compute every possible trip. Unfortunately the 2018 college football season would be over by then, and even worse, humanity isn't even expected to survive nearly that long. Therefore, we need to find a smarter way to optimize this trip if we want a chance of going on it before the Human race goes extinct. In my previous posts, I described a genetic algorithm approach to optimizing trips similar to this one. In this post, I'm going to take the same kind of approach: At first, I'll create 500 random trips and evaluate them according to how well they meet rules #1--#3. After that, I'll take the best of the initial trips, make copies of them, and make small random changes to the copies, for example by picking the UCF @ UConn game on August 30 instead of the Wake Forest @ Tulane game. Finally, I'll throw out all of the old trips that I started with. After I repeat this evaluate-rank-copy-modify-delete process 500 times, I'm left with an optimized trip that maximizes the number of Division I FBS teams that we can see in the 2018 season. And this time we only had to wait a few minutes. Optimized college football trip v1.0 After a few minutes of running, my algorithm output the optimized trip below. Although this trip isn't provably optimal, it provides an itinerary that will take us to 48 different stadiums, watch 90 different teams (including 47 of the top 50 Power 5 teams), and witness 6 of the biggest rivalries in college football---all within the restrictive 51-day regular season schedule. I've provided an interactive map of the trip below. Click here to see an interactive version of the map From the Red River Showdown to the classic Army-Navy Game, this trip will have you visiting many of the nation's iconic stadiums and watching some of the best games that college football has to offer. Boise State's Albertsons Stadium is one of the many unique stadiums you'll visit on this trip If you'd like to look up each game, I've provided a table below. In some cases, the teams are playing on a neutral site and the Home/Away designation doesn't apply. Date Location Home team Away team August 25 Aggie Memorial Stadium, Las Cruces, NM New Mexico State Aggies Wyoming Cowboys August 30 Ross-Ade Stadium, West Lafayette, IN Purdue Boilermakers Northwestern Wildcats August 31 Camp Randall Stadium, Madison, WI Wisconsin Badgers Western Kentucky Hilltoppers September 1 Notre Dame Stadium, Notre Dame, IN Notre Dame Fighting Irish Michigan Wolverines September 2 AT&T Stadium, Arlington, TX LSU Tigers Miami Hurricanes September 3 Doak Campbell Stadium, Tallahassee, FL Florida State Seminoles Virginia Tech Hokies September 7 Gerald J. Ford Stadium, Dallas, TX SMU Mustangs TCU Horned Frogs September 8 Kyle Field, College Station, TX Texas A&M Aggies Clemson Tigers September 13 BB&T Field, Winston-salem, NC Wake Forest Demon Deacons Boston College Eagles September 14 Liberty Bowl Memorial Stadium, Memphis, TN Memphis Tigers Georgia State Panthers September 15 Williams-Brice Stadium, Columbia, SC South Carolina Gamecocks Marshall Thundering Herd September 20 Lincoln Financial Field, Philadelphia, PA Temple Owls Tulsa Golden Hurricane September 21 Memorial Stadium, Champaign, IL Illinois Fighting Illini Penn State Nittany Lions September 22 Jordan-Hare Stadium, Auburn, AL Auburn Tigers Arkansas Razorbacks September 27 Hard Rock Stadium, Miami, FL Miami Hurricanes North Carolina Tar Heels September 28 Folsom Field, Boulder, CO Colorado Buffaloes UCLA Bruins September 29 Memorial Stadium, Lawrence, KS Kansas Jayhawks Oklahoma State Cowboys October 4 TDECU Stadium, Houston, TX Houston Cougars Tulsa Golden Hurricane October 5 LaVell Edwards Stadium, Provo, UT BYU Cougars Utah State Aggies October 6 Cotton Bowl, Dallas, TX Oklahoma Sooners Texas Longhorns October 9 Centennial Bank Stadium, Jonesboro, AR Arkansas State Red Wolves Appalachian State Mountaineers October 11 Bobcat Stadium, San Marcos, TX Texas State Bobcats Georgia Southern Eagles October 12 Rice-Eccles Stadium, Salt Lake City, UT Utah Utes Arizona Wildcats October 13 Bobby Dodd Stadium, Atlanta, GA Georgia Tech Yellow Jackets Duke Blue Devils October 18 Sun Devil Stadium, Tempe, AZ Arizona State Sun Devils Stanford Cardinal October 19 Albertsons Stadium, Boise, ID Boise State Broncos Colorado State Rams October 20 Neyland Stadium, Knoxville, TN Tennessee Volunteers Alabama Crimson Tide October 23 Ladd-Peebles Stadium, Mobile, AL South Alabama Jaguars Troy Trojans October 25 Milan Puskar Stadium, Morgantown, WV West Virginia Mountaineers Baylor Bears October 26 TCF Bank Stadium, Minneapolis, MN Minnesota Golden Gophers Indiana Hoosiers October 27 EverBank Field, Jacksonville, FL Georgia Bulldogs Florida Gators October 30 Doyt Perry Stadium, Bowling Green, OH Bowling Green Falcons Kent State Golden Flashes October 31 Glass Bowl, Toledo, OH Toledo Rockets Ball State Cardinals November 1 Spectrum Stadium, Orlando, FL UCF Knights Temple Owls November 2 Scott Stadium, Charlottesville, VA Virginia Cavaliers Pittsburgh Panthers November 3 Ohio Stadium, Columbus, OH Ohio State Buckeyes Nebraska Cornhuskers November 6 UB Stadium, Buffalo, NY Buffalo Bulls Kent State Golden Flashes November 7 Yager Stadium, Oxford, OH Miami (OH) RedHawks Ohio Bobcats November 8 Carter-Finley Stadium, Raleigh, NC NC State Wolfpack Wake Forest Demon Deacons November 9 Carrier Dome, Syracuse, NY Syracuse Orange Louisville Cardinals November 10 Los Angeles Memorial Coliseum, Los Angeles, CA USC Trojans California Golden Bears November 13 Scheumann Stadium, Muncie, IN Ball State Cardinals Western Michigan Broncos November 14 Peden Stadium, Athens, OH Ohio Bobcats Buffalo Bulls November 15 TDECU Stadium, Houston, TX Houston Cougars Tulane Green Wave November 16 Gerald J. Ford Stadium, Dallas, TX SMU Mustangs Memphis Tigers November 17 Bill Snyder Family Stadium, Manhattan, KS Kansas State Wildcats Texas Tech Red Raiders November 20 Waldo Stadium, Kalamazoo, MI Western Michigan Broncos Northern Illinois Huskies November 22 Vaught-Hemingway Stadium, Oxford, MS Ole Miss Rebels Mississippi State Bulldogs November 23 Martin Stadium, Pullman, WA Washington State Cougars Washington Huskies November 24 Spartan Stadium, East Lansing, MI Michigan State Spartans Rutgers Scarlet Knights December 8 Lincoln Financial Field, Philadelphia, PA Navy Midshipmen Army Black Knights This trip is a still a work in progress, and there are a couple areas that I'd like to improve upon. For one, there are several iconic football stadiums that this trip misses out on, The Big House among them. I'd like to find a way to incorporate those stadiums into this trip without missing out on any more teams. I'll make sure to keep this post updated as I make progress. That's all for now. Do you have any ideas on how to improve this trip? Leave your thoughts in the comments. --- ## How Americans make a living based on their age URL: https://www.randalolson.com/2018/03/06/how-americans-make-a-living-based-on-their-age/ Published: 2018-03-06 Categories: analysis, data visualization Tags: data visualization, employment, united states, usa, visualization Randy Olson revisits the U.S. Bureau of Labor Statistics employment data to look at how Americans make a living based on their age. As I discussed in my last post, certain industries in the U.S. draw a younger (or older) demographic than others. There are a variety of reasons that younger people are more likely to work in a shoe store than a funeral home, but I'm not going to touch on that here. Instead, in this post I want to provide a more detailed view of the U.S. Bureau of Labor Statistics employment data. In the chart below, I plotted employment in the 13 high-level industries tracked by the BLS as a function of age groups ranging from 16-19 to 65+. To read this chart, start with the x-axis, which shows the age groups tracked by the BLS in order. The y-axis shows the percent of Americans in that age group who are employed in each industry. Lastly, the color of the lines indicate the industry, which are listed on the side of the cart. For example, the light blue line sitting at 25% for the 16-19 age group indicates that 25% of 16-19 year olds work in the "Wholesale and retail trade." You can follow that light blue line to see how different age groups move in or out of that industry. Perhaps the most noticeable data point on this chart is that roughly 40% of all employed 16-19 year olds work in the leisure and hospitality industry, primarily in the food service industry. The leisure and hospitality industry quickly loses popularity with 20+ year olds, and constitutes less than 10% of the jobs that 30+ year olds hold. Some other notables: The wholesale and retail trades are similarly popular among 16-19 year olds but less so among 20+ year olds. However, it seems to resurge in popularity among 55+ year olds, possibly due to retirement from other professions. Manufacturing jobs see a major drop among 65+ year olds, likely due to retirement from those physically demanding jobs. Construction jobs follow a similar a similar path. Education and health service as well as professional and business service jobs grow in popularity as workers grow older, likely due to the educational and experience requirements of those jobs. Both industries seem fairly resilient to retirement age, unlike manufacturing and construction. You can download the dataset as a CSV file here if you'd like to work with it yourself. --- ## Top 10 oldest and youngest industries in the U.S. URL: https://www.randalolson.com/2018/03/05/top-10-oldest-and-youngest-industries-in-the-u-s/ Published: 2018-03-05 Categories: analysis, data visualization Tags: data visualization, employment, united states Ever wondered what industries employ the youngest and oldest people? Wonder no more. Randy Olson analyzes the Bureau of Labor Statistics' employment data to discover the oldest and youngest industries in the U.S. Ever wondered what industries employ the youngest and oldest people? Wonder no more. To create the visualization below, I pulled the employment data from the U.S. Bureau of Labor Statistics web site to rank all of the industries by the median age of the people working in them. Perhaps unsurprisingly, most of the youngest industries are in retail and service: retail store clerks, restaurant waiters, and bartenders are a few examples. The most surprising finding for me here was how young shoe salespeople are on average, coming in at an average of 25.6 years old---more than 16 years younger than the national average of all industries at 42.2 years old. If you visit your local shoe store, there's a good chance that everyone working there will be under 30. Times sure have changed since Al Bundy ruled the shoe market! On the other end of the spectrum, the oldest industries are quite varied, ranging from millworkers to bus drivers to repairmen. There is some irony that funeral home workers host the oldest employees on average, given the focus of their occupation. For those wondering, fuel dealers differ from gas stations in that they sell and deliver the fuel themselves, typically in large quantities. If you want to look up your occupation, use the BLS's employment records. Did you know that the average programmer in the U.S. is 41 years old? --- ## Machine Learning Madden NFL: The best player position switches for Madden 17 URL: https://www.randalolson.com/2017/01/20/machine-learning-madden-nfl-the-best-player-position-switches-for-madden-17/ Published: 2017-01-20 Categories: analysis, machine learning Tags: football, machine learning, madden, nfl Randy Olson uses machine learning to figure out the best player position switches in Madden 17. A couple weeks ago, I wrote about my initial efforts toward using machine learning to model the "master equations" that govern the Madden NFL player ratings system. This week, I'd like to put those models to use to compute the best player position switches for Madden 17. Sometimes NFL players are assigned to play a particular position for their team even though they're better suited to play at another position. Perhaps their team had a glaring weakness at that position, or there was already a better player filling their preferred position. In Madden 17, these position mismatches cause some players to have an embarrassingly low overall player rating (OVR) even though they're star players waiting for their time to shine. I'm going to highlight a handful of those players below, but I've also uploaded a spreadsheet of Madden 17 player ratings at all positions so you can sort through them yourself. The >95 OVR club Perhaps the worst position mismatch in the Madden 17 preseason teams is Khalil Mack: Image courtesy of Wikipedia Technically, Khalil Mack should've been a part of Madden 17's coveted 99 OVR club, but since he started the season at left outside linebacker (LOLB), he was assigned a 94 OVR. A wise Madden coach will quickly move Mack over to the left end (LE) position, where Madden will rate him at 99 OVR and Mack will rain sacks on QBs all season. Perhaps it makes sense to put Mack at LOLB in a 3-4 defense, but don't put his pass rushing talents to waste as a linebacker in a 3-4 defense. Here's all of the Madden 17 players that could've made the >95 OVR club: Player Team Position OVR Recommended Position New OVR Khalil Mack Raiders LOLB 94 LE 99 (+5) Derrick Johnson Chiefs MLB 90 SS 98 (+8) Pernell McPhee Bears LOLB 92 DT 98 (+6) Brandon Marshall Broncos MLB 90 SS 96 (+6) Sean Lee Cowboys ROLB 89 SS 96 (+7) Thomas Davis Sr. Panthers LOLB 88 SS 96 (+8) Biggest gainers Below is a list of players who improve in OVR the most simply by changing positions. Notably, Lee Smith can jump from a mediocre Madden tight end (TE) to an above average fullback (FB) thanks to his exceptional blocking skills. Player Team Position OVR Recommended Position New OVR Zak DeOssie Giants TE 55 ROLB 75 (+20) Lee Smith Raiders TE 75 FB 87 (+12) Christian Kirksey Browns MLB 74 SS 85 (+11) Jonathan Anderson Bears MLB 66 SS 77 (+11) Cordarrelle Patterson Vikings WR 70 HB 81 (+11) Deone Bucannon Cardinals MLB 78 SS 88 (+10) Best out-of-position kickers And just for fun---and because the NFL had a bit of a kicker SNAFU this year---below is a list of the non-kickers who could become the best kickers if necessity called. Funny enough, according to Madden, Ryan Tannehill (the Dolphins' starting QB) would make almost as good a kicker as he would a QB. Player Team Position OVR Recommended Position New OVR Ryan Tannehill Dolphins QB 79 K 76 Tim Hightower Saints HB 73 K 75 Odell Beckham Jr. Giants WR 93 K 74 Mohamed Sanu Sr. Falcons WR 77 K 67 Dontrelle Inman Chargers WR 72 K 67 Ben Roethlisberger Steelers QB 93 K 66 Feel free to look through the full spreadsheet and report your findings in the comments. There's far more useful position switches than what I reported here. --- ## Machine Learning Madden NFL: How Madden player ratings are actually calculated URL: https://www.randalolson.com/2017/01/10/machine-learning-madden-nfl-how-madden-player-ratings-are-actually-calculated/ Published: 2017-01-10 Categories: machine learning Tags: football, machine learning, madden, nfl Randy Olson uses machine learning to learn how Madden calculates a player's overall rating. For the past few months, I've been playing Madden NFL 17 in my free time. I really enjoy the team-building aspect of the franchise mode, where I've taken on challenges such as finally bringing the Lombardi trophy home to Philadelphia. As a data nerd, one thing I love about Madden NFL is their player rating system, where they assign numerical values to practically every aspect of every player's abilities. These ratings are then summarized into an overall player rating (OVR for short), which is a fairly good indication of how skilled a player is and how dominant they will be on the football field. Image courtesy of EA Sports A couple years ago, FiveThirtyEight published a great data-driven feature explaining how Madden NFL's player ratings are made, and how essentially one person is behind the whole system. Perhaps most useful for Madden NFL players, they provided a chart showing the most important traits for each position in the game. A lot has changed in the past couple years, though, and for the longest time I've wanted to figure out the "master equation" that actually turns player ratings into their overall rating so I know what traits to focus on developing. That's the goal of this post, where I'm going to use machine learning to attempt to discover that "master equation." If you don't care about how I computed the ratings, feel free to skip ahead to the "master equation" section. Analyzing Madden NFL player ratings The Madden NFL folks have been pretty awesome about sharing their player rating data. We can look up and download all of the active player ratings directly from the Madden NFL web site, which I've done and shared in a Google doc here. For this post, I'm working with the preseason ratings. Since no analysis is complete without some basic exploratory data analysis, let's take a look at the distribution of overall player ratings, which is what we're looking to model. The chart below is a violin plot, which is a combination of a box plot and histogram. Basically, it shows the distribution of Madden player OVRs. Madden player OVRs range from 40 (worst) to 99 (best), with a majority of the players falling between 65 and 75 OVR. Only a handful of players fall into the coveted 99 rating spot, which include current NFL greats such as Patriots TE Rob Gronkowski (when he's healthy) and Broncos linebacker Von Miller (the highest paid defensive player in NFL history). There's an interesting blip of players in the low-40s OVR range, which are better highlighted when we break down the OVRs by position: All of the players rated in the low 40s OVR are TEs, which seems quite odd. Do NFL coaches like to hoard crappy TEs? Why would professional football teams keep players who are rated so poorly that they're statistical anomalies (in a bad way) when compared to the rest of the NFL? After a little digging into the player rating spreadsheet, it seems that nearly all of these low-40s OVR TEs are actually long snappers, and the Madden NFL folks decided to make them TEs for some reason. Since Madden NFL doesn't currently feature long snappers nor a long snapping skill, my guess is that most of these long snappers don't even make the preseason cuts in most Madden games. Another interesting finding from the above chart is that kickers and punters tend to have the highest OVR, on average. This is likely because there's so few kickers and punters in the league (typically, only one of each per team), and few teams (except for the 2016 Buccaneers) keep a kicker or punter around if they don't excel at their job. There are many more interesting statistics to look at in this data set, but in the interest of brevity, I'll leave those charts for another post. Finding the master equations Now that we have a basic understanding of the data set, we want to find the "master equation" that turns player skill ratings into their OVR, which represents how good of a player they are in Madden. To accomplish this goal, I trained a linear model (with Lasso regularization) to predict every player's OVR based on their skill ratings. Essentially, I gave the model every player's skill ratings (speed, strength, throw power, etc.) and asked it to predict each player's OVR for me. Since each position's OVR is calculated differently---throw power is far more important for quarterbacks than defensive tackles, for example---I grouped the players by their position and trained a different linear model for each position. Thus, we're no longer looking for a master equation; we're looking for the master equations. One of the main advantages of linear models is their interpretability. Once a linear model is trained on a data set, we can inspect the model, which is simply a sum of the player skill ratings with weights applied to them. For example, maybe the QB master equation would look like: QB OVR = 10 + 0.5 x Throw Power + 0.25 x Throw Accuracy Short + 0.2 x Throw Accuracy Medium + 0.1 x Throw Accuracy Deep So if we want to calculate Tom Brady's OVR, we would plug his ratings into the equation: Tom Brady OVR = 10 + 0.5 x 94 + 0.25 x 98 + 0.2 x 98 + 0.1 x 84 = ... and so on. Obviously the above equation was made up, but let's take a look at the master equations that the linear models discovered. The master equations Below, I made a heatmap of the player traits that affect each position's OVR. Darker purple cells indicate important traits for the position, whereas light purple or blank traits have little to no effect. The number in each cell is the linear model's coefficient, which is the number that's multiplied by the trait in the master equation. Here are the intercepts for each linear model: Position Intercept QB -67.83 HB -62.34 FB -76.59 WR -52.84 TE -61.99 LT -43.45 LG -55.1 C -56.83 RG -52.97 RT -44.89 LE -69.02 DT -61.49 RE -63.67 LOLB -49.92 MLB -56.68 ROLB -50.14 CB -54.81 SS -47.58 FS -45.75 K 0.72 P 1.64 The intercept is the number you add at the beginning of the equation. For the simplest example, to calculate a Punter's OVR: Punter OVR = 1.64 + 0.39 x Awareness + 0.32 x Kick Power + 0.27 x Kick Accuracy What's more interesting, though, are the most important traits for each position: Awareness (AWR) dominates as the most important trait for nearly every position, even though many Madden players don't think Awareness has much of an effect in Madden Speed (SPD) has a relatively minor impact on OVR for even the most speedy positions (HB, WR, CB), even though Speed is widely considered the most important trait by Madden players Play Recognition (PRC) trumps Awareness (AWR) for Strong Safeties Impact Blocking (IMP) is far more important for Fullbacks than the original FiveThirtyEight study suggested ... which seems to suggest that we could make our own master equations to find every Madden player's "True OVR." We'll leave that task for a future post, though. Caveats No machine learning analysis comes without its assumptions and caveats, and I'll list the caveats of this analysis here: The estimates from these equations are just that---estimates. The predicted OVR may be a point or two off from the real OVR, but most of the time it should be correct. The team's scheme (4-3 vs 3-4 defense, for example) will affect a player's OVR for that team, which may cause some of the predicted OVRs to be off by a point or two. These equations are based on the Madden 17 preseason ratings. They may change in the future at the Madden developer's whims. These equations are from a linear model, and it's possible that some traits have diminishing returns at some point. If some traits have diminishing returns, we may need to build a more complex model for those traits. Feel free to add your thoughts in the comments. Now that we have the Madden master equations, what can we do with them? --- ## Python 2.7 still reigns supreme in pip installs URL: https://www.randalolson.com/2016/09/03/python-2-7-still-reigns-supreme-in-pip-installs/ Published: 2016-09-03 Categories: data visualization, python Tags: ipython, pandas, python Randy Olson analyzes trends in Python package installs to observe how the transition from Python 2 to 3 is progressing The Python 2 vs. Python 3 divide has long been a thorn in the Python community's side. On one hand, Python package developers face the challenge of supporting two incompatible versions of Python, which is time that could be better spent improving the package. On the other hand, many Python users are reluctant to upgrade from Python 2 to 3 because of the time commitment such an upgrade entails. The Python Software Foundation's official stance on the matter is: Python 2.x is legacy, Python 3.x is the present and future of the language (Upgrading Python 2 code to Python 3 isn't that bad, by the way.) I like to check in on the Python community's transition from 2 to 3 every once in a while, so I figured it was about time for another check. Conveniently, Juan Pablo posted a preliminary analysis yesterday looking at the evolution of pip package downloads by Python version over the past couple months. Below, I will delve into that data a bit more to see what insights we can draw. Overall Python downloads If we look at all pip installs in July and August '16, Python 2.7 still comprises roughly 90% of all pip installs. There are some interesting day-to-day fluctuations that presumably show fewer users installing Python packages on the weekend, but overall there are about 10,000,000 packages installed on Python 2.7 distributions every day. Of course, the statistics above capture all pip installs from all Python packages. What about the Scientific Python (SciPy) stack, which most of my readers are probably concerned about? SciPy stack downloads To look at Python usage across the SciPy stack, I used the same query as above except I limited the search to the following packages: NumPy SciPy matplotlib pandas SymPy IPython Jupyter nose scikit-learn scikit-image Seaborn Bokeh (I added the last two per my personal opinion; the others are referenced on the SciPy page.) Much to my dismay, even in the SciPy stack pip is used to install an order of magnitude more packages on Python 2.7 distributions than Python 3 (roughly 80% of all installs), with little sign of slowing down. At this point, we have to wonder: will Python's scientific Python community be ready when support for Python 2 is officially dropped in 2020? Breakdown by packages in the SciPy stack I was also curious to see the Python usage statistics broken down by the various scientific Python packages, so that's what I've plotted below. In general, it seems that it's important for the scientific Python packages to primarily support Python 2.7, 3.4, and 3.5, with a handful of packages even having a large Python 2.6 user base. Conclusions Python 2 still seems to be the most-used version of Python by far, at least in terms of packages installed via pip. With Python 2's end of life in the near future, it's time for the Python community to start having a serious conversation about how to smooth the transition from Python 2 to 3. Aside from 2to3 for automatic code translation, most major packages providing Python 3 support, and several guides focused on porting Python code from 2 to 3, what are we missing? If you're still using Python 2, what would convince you to make the switch? Data source I queried the PSF downloads table on BigQuery using the following query. SELECT CONCAT( DATE(timestamp), '_', REGEXP_EXTRACT(details.python, r'^([2-3].[0-9]).') ) AS date_python, COUNT(details.python) AS downloads FROM (TABLE_DATE_RANGE([the-psf:pypi.downloads], TIMESTAMP('2016-06-01'), TIMESTAMP('2016-09-01'))) WHERE LOWER(details.installer.name) LIKE 'pip' AND (LOWER(file.project) LIKE 'numpy' OR LOWER(file.project) LIKE 'scipy' OR LOWER(file.project) LIKE 'matplotlib' OR LOWER(file.project) LIKE 'pandas' OR LOWER(file.project) LIKE 'sympy' OR LOWER(file.project) LIKE 'ipython' OR LOWER(file.project) LIKE 'jupyter' OR LOWER(file.project) LIKE 'nose' OR LOWER(file.project) LIKE 'scikit-%' OR LOWER(file.project) LIKE 'seaborn' OR LOWER(file.project) LIKE 'bokeh') GROUP BY date_python ORDER BY date_python This query was modified from Juan Pablo's earlier query. --- ## Republican-leaning states tend to have more traffic deaths URL: https://www.randalolson.com/2016/09/03/republican-leaning-states-tend-to-have-more-traffic-deaths/ Published: 2016-09-03 Categories: data visualization Tags: democrat, politics, republican, statistics Randy Olson explores an odd statistic about Republican-leaning states. Back in 2014, the U.S. Department of Transportation released a report on the (normalized) number of traffic deaths in each U.S. state. As I looked through the list, I noticed an odd correlation between the political leanings of a state and its traffic fatalities. To confirm this notion, I used box plots to show the distribution of traffic deaths based on how Republican-leaning each state is, where each dot represents a state. In case you're wondering, Delaware is the Democratic outlier with 1.26 deaths per 100 million miles traveled. Outliers aside, we see an abundantly clear trend: Republican-leaning states tend to have more traffic deaths. To provide another view, I plotted the same data set as a scatter plot and fit a linear regression below. Each dot represents a state. Again, we see a strong relationship between Republican-leaning states and more traffic deaths. At this point, I feel obligated to make the caveat that these charts only establish a correlation, and it doesn't mean that residents of Republican-leaning states experience higher traffic deaths because they tend to vote Republican. However, it sure makes for an interesting statistic to ponder: why would there be a correlation between traffic death rate and tendency to vote Republican? Thanks to David Fairlie for sharing these interesting statistics on /r/DataIsBeautiful. --- ## Evolution of active categorical image classification via saccadic eye movement URL: https://www.randalolson.com/2016/08/13/evolution-of-active-categorical-image-classification-via-saccadic-eye-movement/ Published: 2016-08-13 Categories: machine learning, research Tags: evolutionary computation, image classification, machine learning Randy Olson presents demo videos showing a new approach to image classification, where the machine scans and classifies the image the same way humans would. I put together a couple demo videos for our Active Categorical Classifier (ACC) project that we'll be presenting at the PPSN 2016 conference. If you're interested in this project and can't wait for PPSN, we have: a preprint of the paper on arXiv, the C++ code on GitHub, and a demo of the new Python version on GitHub. In the first video, I show the entire image as the ACC roams around it trying to classify the object. Even though we can see the entire image at once, the ACC only sees the 9 pixels in its immediate vicinity. In the second video below, I only show the pixels that the ACC has seen during the simulation. Can you guess what digit it is before the ACC? I think the second video does a great job of showing how little information the ACC is processing to classify each image. Talk about a light-weight classifier! --- ## The Optimal U.S. National Parks Centennial Road Trip URL: https://www.randalolson.com/2016/07/30/the-optimal-u-s-national-parks-centennial-road-trip/ Published: 2016-07-30 Categories: data visualization, machine learning Tags: national parks, machine learning, optimization, road trip, traveling salesman problem Randy Olson shows how to create an optimized road trip that visits the U.S. National Parks. In August 2016, the National Park Service celebrates their 100th year of managing the United States' system of beautiful national parks. So what's a better way to celebrate 100 years of stewardship than to visit all of the national parks in one epic road trip? If you've followed my blog for the past year or so, you'll know that I've made a hobby of optimizing various road trips around the U.S., so I couldn't pass up on this opportunity to optimize yet another road trip. U.S. National Parks If you're unfamiliar with the U.S. national park system, it consists of 59 protected areas across the U.S. that are managed by the U.S. National Park Service. Many of the national parks are known for their natural beauty, unique geological features, unusual ecosystems, and/or recreational opportunities, which makes them ideal spots to visit if you need a break from the hustle and bustle of the big city. Unique combinations of geologic color and erosional forms decorate a canyon that is 277 river miles (446km) long, up to 18 miles (29km) wide, and a mile (1.6km) deep. Grand Canyon overwhelms our senses through its immense size. [NPS] Ridge upon ridge of forest straddles the border between North Carolina and Tennessee in Great Smoky Mountains National Park. World renowned for its diversity of plant and animal life, the beauty of its ancient mountains, and the quality of its remnants of Southern Appalachian mountain culture, this is America's most visited national park. [NPS] Visit Yellowstone and experience the world's first national park. Marvel at a volcano's hidden power rising up in colorful hot springs, mudpots, and geysers. Explore mountains, forests, and lakes to watch wildlife and witness the drama of the natural world unfold. Discover the history that led to the conservation of our national treasures "for the benefit and enjoyment of the people." [NPS] 12 of the national parks are in Alaska, Hawaii, and other U.S. territories, which can make them difficult to drive to unless you have a flying car. Thus for road trip, we're going to focus on the national parks that span the 48 contiguous states in the mainland United States. Don't worry: that limitation still leaves us 47 national parks, which should be plenty for one road trip. The optimal road trip to the U.S. National Parks In total, this road trip spans 14,498 miles (23,333 km) of road and will take roughly 2 months if you're traveling at a breakneck pace. I've designed this road trip to form a circle around the U.S., so you can hop on at any point and proceed whatever direction you like. Just make sure to follow the agenda from that point on if you want to follow the optimal route! Here's the Google Maps for the full trip: [1] [2] [3] [4] [5] Here's the full list of national parks in order: Grand Canyon National Park, Arizona Petrified Forest National Park, Arizona Saguaro National Park, Arizona Guadalupe Mountains National Park, Texas Carlsbad Caverns National Park, New Mexico Big Bend National Park, Texas Hot Springs National Park, Arkansas Mammoth Cave National Park, Kentucky Great Smoky Mountains National Park, Tennessee Everglades National Park, Florida Dry Tortugas National Park, Florida Biscayne National Park, Florida Congaree National Park, South Carolina Shenandoah National Park, Virginia Acadia National Park, Maine Cuyahoga Valley National Park, Ohio Isle Royale National Park, Michigan Voyageurs National Park, Minnesota Theodore Roosevelt National Park, North Dakota Badlands National Park, South Dakota Wind Cave National Park, South Dakota Rocky Mountain National Park, Colorado Great Sand Dunes National Park and Preserve, Colorado Black Canyon of the Gunnison National Park, Colorado Mesa Verde National Park, Colorado Canyonlands National Park, Utah Arches National Park, Utah Capitol Reef National Park, Utah Bryce Canyon National Park, Utah Zion National Park, Utah Great Basin National Park, Nevada Grand Teton National Park, Wyoming Yellowstone National Park, Wyoming Glacier National Park, Montana North Cascades National Park, Washington Mount Rainier National Park, Washington Olympic National Park, Washington Crater Lake National Park, Oregon Redwood National and State Parks, California Lassen Volcanic National Park, California Yosemite National Park, California Kings Canyon National Park, California Sequoia National Park, California Pinnacles National Park, California Channel Islands National Park, California Joshua Tree National Park, California Death Valley National Park, California Want to make your own road trip? If you like the idea of taking an optimal road trip but don't like the locations I chose, don't fret: you can make your own road trip! This time around, I used the Gurobi TSP solver to optimize this road trip. Check out Nathan Brixius' blog post to learn how to make your own, or check out my alternative methods to optimizing road trips using Python and Google Maps. If Python coding is beyond you, there are web sites like RouteXL.com that will do it for you. They optimize road trips with up to 20 stops for free, and 20+ stops for a nominal fee. Happy road tripping! --- ## Computing optimal road trips on a limited budget URL: https://www.randalolson.com/2016/06/05/computing-optimal-road-trips-on-a-limited-budget/ Published: 2016-06-05 Categories: data visualization, machine learning, python Tags: genetic algorithm, machine learning, optimization, road trip, travel, traveling salesman problem Randy Olson shows you how to compute an epic road trip across the U.S. when you're on a limited budget. About a year ago, I wrote an article introducing the concept of optimizing road trips using a combination of genetic algorithms and Google Maps. During that time, I've given some thought to how I could make that algorithm more useful to folks looking to plan their summer road trips. One thought that struck me was that the road trips I created before were quite grandiose---spanning entire states or even most of Europe---such that only people who had some savings and were able to take a month off of work could even hope to go on one of the trips. In reality, most of us have budgetary constraints on our road trips: we can only spend so much money, or we only have so much time off before we have to get back to work. In this article, I'm going to expand on the idea of optimizing road trips by introducing multi-objective Pareto optimization to the algorithm. I'll briefly describe how Pareto optimization works, and how it helps us optimize road trips on a limited budget. Note: If you're not interested in the technical details of the project, skip down to the 48 U.S. state capitols in 8 1/2 days section. Planning the road trip: U.S. state capitols For this road trip, there is one goal: to take a picture at as many U.S. state capitols as possible. (Bonus points for entertaining or themed pictures!) We will travel only by car, so that rules out Alaska (too far away) and Hawaii (requires a plane flight) and leaves us with the 48 contiguous states (excluding D.C.). Whenever possible, we will avoid routes that require us to travel through foreign countries, as entering/leaving the country requires a passport and border control tends to slow things down. Lastly, to clarify: The goal is to visit the capitol buildings, not the city the buildings are located in (i.e., state capitals). That said, by going on this road trip we're in for an epic journey and some beautiful architecture. Image credit: Daniel Mennerich Recap: Optimizing road trips With the list of U.S. state capitols in hand, the next step is to find the "true" distance between all of the capitols by car. Since we can't just drive a straight line between every capitol---driving by car has this pesky limitation of having to stay on roads---we needed to find the shortest route by road between every capitol. If you've ever used Google Maps to get the directions between two addresses, that's basically what we have to do here. Except this time, we need to look up 2,256 directions to get the "true" distance between all 48 state capitols---a monumental task if we have to do it by hand. Thankfully, the Google Maps API makes this information freely available, so all it takes is a short Python script to calculate the distance and time driven for all 2,256 routes between the 48 capitols. Now with the 2,256 capitol-capitol distances, our next step is to approach the task as a traveling salesman problem: We need to order the list of capitols such that the total distance traveled between them is as small as possible if we visited them in order. This means finding the route that backtracks as little as possible, which is especially difficult when visiting Florida and the Northeast. If you've read my Where's Waldo? article, you're already aware of how difficult it can be to solve route optimization problems like this one. With 48 landmarks to put in order, we would have to exhaustively evaluate 1.24 x 1061 possible routes to find the shortest one. To provide some context: If you started computing this problem on your home computer right now, you'd find the optimal route in about 3.98 x 1049 years---long after the Sun has entered its red giant phase and devoured the Earth. This complication is why Google Map's route optimization service only optimizes routes of up 10 waypoints, and the best free route optimization service only optimizes 20 waypoints unless you pay them a lot of money to dedicate some bigger computers to it. The traveling salesman problem is so notoriously difficult to solve that even xkcd poked fun at it: Clearly, we need a smarter solution if we want to take this road trip in our lifetime. Thankfully, the traveling salesman problem has been well-studied over the years and there are many ways for us to solve it in a reasonable amount of time. If we're willing to accept that we don't need the absolute best route between all of the capitols, then we can turn to smarter techniques such as genetic algorithms to find a solution that's good enough for our purposes. Instead of exhaustively looking at every possible solution, genetic algorithms start with a handful of random solutions and continually tinkers with these solutions---always trying something slightly different from the current solutions and keeping the best ones---until they can't find a better solution any more. I've included a visualization of a genetic algorithm optimizing one road trip below. Multi-objective Pareto optimization Normally in optimization problems, we want to maximize or minimize one criteria: Maximize how much money we make, or minimize the chance of an accident occurring. In multi-objective Pareto optimization, we can simultaneously optimize many criteria---for example, maximizing how many states we visit, while at the same time minimizing the total time we spend driving for the road trip. In the chart below, each dot corresponds to one road trip. Watch as the genetic algorithm simultaneously optimizes 48 road trips. What's particularly useful about Pareto optimization is that at the end of the optimization process, we have a Pareto front to choose from that lists the trade-offs between what we're trying to optimize. In the above chart, we see that the more states we visit, the longer the trip will take. If we only have 2 days to take a trip, for example, the Pareto front provides a plethora of trips to choose from: Maybe we should visit only a few capitols and have plenty of time to explore them, or perhaps we should visit as many capitols as possible in 2 days. The choice is ours. In the animated map below, I've visualized all 48 of the optimized routes from the Pareto optimization process. Notice how each route differs slightly, for example, the optimized route that reaches 7 capitols is fairly different from the optimized route that reaches 8 capitols. 48 U.S. state capitols in 8 1/2 days After running on my laptop for about 20 minutes, the genetic algorithm reached an optimized solution that makes a complete trip to all of the U.S. state capitols in only 13,310 miles (21,420 km) of driving. I've mapped that route below. Click here for an interactive version of the map Assuming no traffic, this road trip will take about 8 1/2 days of driving in total, so you better bring a big water bottle. The best part is that this road trip is designed so that you can start anywhere on the route: As long as you follow the route from wherever you start, you'll hit every state capitol in the 48 contiguous U.S. states, and as an added bonus, you can even add Washington, D.C. to the route without adding any extra miles. Here's the full list of capitols in order: State House, 107 North Main Street, Concord, NH 03303 Maine State House, Augusta, ME 04330 Vermont State House, 115 State Street, Montpelier, VT 05633 New York State Capitol, State St. and Washington Ave, Albany, NY 12224 New Jersey State House, Trenton, NJ 08608 Pennsylvania State Capitol Building, North 3rd Street, Harrisburg, PA 17120 West Virginia State Capitol, Charleston, WV 25317 Ohio State Capitol, 1 Capitol Square, Columbus, OH 43215 Kentucky State Capitol Building, 700 Capitol Avenue, Frankfort, KY 40601 Tennessee State Capitol, 600 Charlotte Avenue, Nashville, TN 37243 Indiana State Capitol, Indianapolis, IN 46204 Michigan State Capitol, Lansing, MI 48933 Illinois State Capitol, Springfield, IL 62756 2 E Main St, Madison, WI 53703 Minnesota State Capitol, St Paul, MN 55155 500 E Capitol Ave, Pierre, SD 57501 North Dakota State Capitol, Bismarck, ND 58501 Montana State Capitol, 1301 E 6th Ave, Helena, MT 59601 Washington State Capitol Bldg, 416 Sid Snyder Ave SW, Olympia, WA 98504 Oregon State Capitol, 900 Court St NE, Salem, OR 97301 L St & 10th St, Sacramento, CA 95814 Nevada State Capitol, Carson City, NV 89701 700 W Jefferson St, Boise, ID 83720 Utah State Capitol, Salt Lake City, UT 84103 Wyoming State Capitol, Cheyenne, WY 82001 200 E Colfax Ave, Denver, CO 80203 New Mexico State Capitol, Santa Fe, NM 87501 Arizona State Capitol, 1700 W Washington St, Phoenix, AZ 85007 Texas Capitol, 1100 Congress Avenue, Austin, TX 78701 Oklahoma State Capitol, Oklahoma City, OK 73105 300 SW 10th Ave, Topeka, KS 66612 Nebraska State Capitol, 1445 K Street, Lincoln, NE 68509 Iowa State Capitol, 1007 E Grand Ave, Des Moines, IA 50319 Missouri State Capitol, Jefferson City, MO 65101 Arkansas State Capitol, 500 Woodlane Street, Little Rock, AR 72201 400-498 N West St, Jackson, MS 39201 Louisiana State Capitol, Baton Rouge, LA 70802 402 S Monroe St, Tallahassee, FL 32301 Alabama State Capitol, 600 Dexter Avenue, Montgomery, AL 36130 Georgia State Capitol, Atlanta, GA 30334 South Carolina State House, 1100 Gervais Street, Columbia, SC 29201 North Carolina State Capitol, Raleigh, NC 27601 Virginia State Capitol, Richmond, VA 23219 Maryland State House, 100 State Cir, Annapolis, MD 21401 Legislative Hall: The State Capitol, Legislative Avenue, Dover, DE 19901 Connecticut State Capitol, 210 Capitol Ave, Hartford, CT 06106 Rhode Island State House, 82 Smith Street, Providence, RI 02903 Massachusetts State House, Boston, MA 02108 Here's the Google Maps for the road trip: [1] [2] [3] [4] [5] (Note that Google Maps itself only allows 10 waypoints to be routed at a time, hence why there's multiple Maps links.) 10 U.S. state capitols in 24 hours At this point, some of you might be scratching your heads. "Didn't you promise to stop showing us grandiose road trips, Randy?", I imagine you thinking. This is where the Pareto optimization aspect comes in: If we look at the final Pareto front, we don't have to pick the 48-state road trip. We have a whole range of road trips to choose from. Let's say, for example, that we only have 24 hours to dedicate to the road trip. If that's the case, then we can look at the Pareto front above and see that we can reach 10 state capitols and arrive back home in less than 24 hours. That's pretty amazing to think that we can leave one morning, visit 10 state's capitols, and be back in time for breakfast the next day. As you might expect, this road trip is in the Northeastern U.S.: Here's the Google Map for the 24-hour road trip: [1] If you want to look up the other road trips from the Pareto front, I've uploaded them on GitHub. In theory, we can expand this idea to all kinds of budgets. If we only have $100 to spend on gas, we can add gas costs to the Pareto front. If we can only average $50/night on the hotel, we can add the average hotel cost at each stop to the Pareto front. And so on. At this point, the only limit is your imagination. "This is awesome! How can I optimize my own road trip?" If you were inspired by this article and want to make your own road trip, I've released the Python code I used in this project with an open source license and instructions for how to optimize your custom road trip. You can find the code here. You should also check out Nathan Brixius' solution to this challenge using a technique from operations research. Nathan was kind enough to share all of his Python code as well. Note: I don't make custom road trips upon request; I simply don't have the time. However, if you have a neat road trip idea that might be interesting to many people, please feel free to email me your idea. Conclusions I'm reminded of the quote: The world is a book, and those who do not travel read only one page. I hope that this article---through its odd mixture of travel, machine learning, and visualization---has inspired you to go out and embark on your own road trip. Whether it's a trip that a computer optimized in a few minutes or a trip that took you several weeks to hand-design, it only matters that you travel and experience the world from a fresh perspective. Happy road tripping! --- ## How long does the average man last in bed? URL: https://www.randalolson.com/2016/05/31/how-long-does-the-average-man-last-in-bed/ Published: 2016-05-31 Categories: data visualization Tags: sex, survey Randy Olson turns to the data to shed light on how long men last in bed. A couple weeks ago, I ran across a research paper that inadvertently answered one of those awkward questions that so few people get the chance to talk about: How long does the average man last in bed? I was particularly interested in this study because men's prowess in the bedroom is often exaggerated, and this study provided an opportunity to shed some light on the truth. The study is (sadly) locked behind a paywall, so I'll do my best to reproduce the results here. You see, these researchers were actually studying premature ejaculation. To do so, they recruited about 500 random heterosexual couples from The Netherlands, Spain, Turkey, the UK, and the US and asked them to track the duration of their sexual encounters using a timer and diary. The researchers asked the couples to start the timer as soon as vaginal penetration occurred, and stop the timer as soon as the man ejaculated. Quite romantic, as you might imagine. After 4 weeks of stop watches and sex diaries, the researchers followed up with the couples to collect the diaries and ask them various demographic questions. Below, I've plotted the distribution of how long each man lasted before orgasming (averaged over all his encounters). To provide a more succinct view of the data, I also created a box plot to summarize the distribution: The average (median) time before orgasm was about 6 minutes, and ranged from a blissful 6 seconds to a marathon-paced 53 minutes. The majority of men lasted between 4 to 11 minutes, with anyone lasting longer than 21 minutes being considered an outlier. What I found particularly interesting about this study was the men's tendency to overestimate the duration of their sexual encounters: According to the authors, the men's estimates averaged about 1.9 minutes longer than they really were---about a 31% overestimation over the 6-minute average---which really highlights our tendency to overestimate our performance in the bedroom. So, there we have it. Our collective fascination with hours-long romps in the bedroom doesn't really hold up in the data, and most couples are probably quite happy for it. Hopefully studies like this one will help us ground our expectations in reality, rather than trying to live up to fantasy. If you're interested in reading more about the study and don't have institutional access, I hear that a certain science hub might have a copy laying around. --- ## Why is Reddit replacing Imgur? URL: https://www.randalolson.com/2016/05/25/why-is-reddit-replacing-imgur/ Published: 2016-05-25 Categories: data visualization, reddit Tags: Imgur, reddit, top posts Randy Olson looks into the data to find out why Reddit is trying to replace Imgur. In a surprise move this week, Reddit has started rolling out their own in-house image hosting service. This appears to be a direct move to replace the many image hosting services that have sprung up around Reddit---Imgur, in particular. Here's a demo of the new service, provided by Reddit: Although their official statement for this change is to "[bring us] a more seamless experience" on Reddit, I have a feeling there's a more practical reason behind Reddit's change. You see, Imgur started as a humble image hosting service to meet Reddit's image sharing needs in early 2009. Nowadays, well over half of Reddit's popular content links to Imgur, causing Imgur to become one of the top 50 most-visited web sites on the Internet (according to Alexa). (Data from Google BigQuery's Reddit post database; a post is considered "popular" if it achieves a score of >= 100.) Imgur has now grown into a full-fledged online community focused on image sharing, and is arguably a direct competitor to Reddit. In a sense, Imgur has gotten too big for its britches, and it's probably too risky for Reddit to continue relying on a direct competitor for image hosting. Combine these observations with Imgur's recent aggressive monetization campaign---link-jacking direct image links on mobile devices to display ads, for example---and it's really no surprise that Reddit wants an in-house image hosting service to supplant Imgur. Until Reddit makes an official statement, all we can do is speculate... but this explanation makes sense to me. On the bright side, Reddit seems to be growing beyond Imgur. While the number of popular Imgur links seems to have stagnated in 2015, we're seeing more and more popular Reddit posts including self-posts, links to other sites, and even other image hosts. We'll have to check back in a few months to see how well Reddit's new image hosting service takes off. What do you think? Is the writing on the wall for Imgur? --- ## TPOT: A Python tool for automating data science URL: https://www.randalolson.com/2016/05/08/tpot-a-python-tool-for-automating-data-science/ Published: 2016-05-08 Categories: machine learning, python, research Tags: automation, data science, machine learning Randy Olson demonstrates why designing machine learning pipelines is difficult, and how it can be automated using TPOT. Machine learning is often touted as: A field of study that gives computers the ability to learn without being explicitly programmed. Despite this common claim, anyone who has worked in the field knows that designing effective machine learning systems is a tedious endeavor, and typically requires considerable experience with machine learning algorithms, expert knowledge of the problem domain, and brute force search to accomplish. Thus, contrary to what machine learning enthusiasts would have us believe, machine learning still requires a considerable amount of explicit programming. In this article, we're going to go over three aspects of machine learning pipeline design that tend to be tedious but nonetheless important. After that, we're going to step through a demo for a tool that intelligently automates the process of machine learning pipeline design, so we can spend our time working on the more interesting aspects of data science. Let's get started. Model hyperparameter tuning is important One of the most tedious parts of machine learning is model hyperparameter tuning. Support vector machines require us to select the ideal kernel, the kernel's parameters, and the penalty parameter C. Artificial neural networks require us to tune the number of hidden layers, number of hidden nodes, and many more hyperparameters. Even random forests require us to tune the number of trees in the ensemble at a minimum. All of these hyperparameters can have significant impacts on how well the model performs. For example, on the MNIST handwritten digit data set: If we fit a random forest classifier with only 10 trees (scikit-learn's default): import pandas as pd import numpy as np from sklearn.ensemble import RandomForestClassifier from sklearn.cross_validation import cross_val_score mnist_data = pd.read_csv('https://raw.githubusercontent.com/rhiever/Data-Analysis-and-Machine-Learning-Projects/master/tpot-demo/mnist.csv.gz', sep='\t', compression='gzip') cv_scores = cross_val_score(RandomForestClassifier(n_estimators=10, n_jobs=-1), X=mnist_data.drop('class', axis=1).values, y=mnist_data.loc[:, 'class'].values, cv=10) print(cv_scores) [ 0.93461813 0.96287836 0.94688749 0.94072275 0.95114286 0.94570653 0.94884253 0.94311848 0.93825043 0.95668954] print(np.mean(cv_scores)) 0.946885709001 The random forest achieves an average of 94.7% cross-validation accuracy on MNIST. However, what if we tuned that hyperparameter a little bit and provided the random forest with 100 trees instead? cv_scores = cross_val_score(RandomForestClassifier(n_estimators=100, n_jobs=-1), X=mnist_data.drop('class', axis=1).values, y=mnist_data.loc[:, 'class'].values, cv=10) print(cv_scores) [ 0.96259814 0.97829812 0.9684466 0.96700471 0.966 0.96399486 0.97113461 0.96755752 0.96397942 0.97684391] print(np.mean(cv_scores)) 0.968585789367 With such a minor change, we improved the random forest's average cross-validation accuracy from 94.7% to 96.9%. This small improvement in accuracy can translate into millions of additional digits classified correctly if we're applying this model on the scale of, say, processing addresses for the U.S. Postal Service. Never use the defaults for your model. Hyperparameter tuning is vitally important for every machine learning project. Model selection is important We all love to think that our favorite model will perform well on every machine learning problem, but different models are better suited for different tasks. For example, if we're working on a signal processing problem where we need to classify whether there's a "hill" or "valley" in the time series: And we apply a "tuned" random forest to the problem: import pandas as pd import numpy as np from sklearn.ensemble import RandomForestClassifier from sklearn.linear_model import LogisticRegression from sklearn.cross_validation import cross_val_score hill_valley_data = pd.read_csv('https://raw.githubusercontent.com/rhiever/Data-Analysis-and-Machine-Learning-Projects/master/tpot-demo/Hill_Valley_without_noise.csv.gz', sep='\t', compression='gzip') cv_scores = cross_val_score(RandomForestClassifier(n_estimators=100, n_jobs=-1), X=hill_valley_data.drop('class', axis=1).values, y=hill_valley_data.loc[:, 'class'].values, cv=10) print(cv_scores) [ 0.64754098 0.64754098 0.57024793 0.61983471 0.62809917 0.61983471 0.70247934 0.59504132 0.49586777 0.65289256] print(np.mean(cv_scores)) 0.617937948787 Then we're going to find that the random forest isn't well-suited for signal processing tasks like this one when it achieves a disappointing average of 61.8% cross-validation accuracy. What if we tried a different model, for example a logistic regression? cv_scores = cross_val_score(LogisticRegression(), X=hill_valley_data.drop('class', axis=1).values, y=hill_valley_data.loc[:, 'class'].values, cv=10) print(cv_scores) [ 1. 1. 1. 0.99173554 1. 0.98347107 1. 0.99173554 1. 1. ] print(np.mean(cv_scores)) 0.996694214876 We'll find that a logistic regression is well-suited for this signal processing task---in fact, it easily achieves near-100% cross-validation accuracy without any hyperparameter tuning at all. Always try out many different machine learning models for every machine learning task that you work on. Trying out---and tuning---different machine learning models is another tedious yet vitally important step of machine learning pipeline design. Feature preprocessing is important As we've seen in the previous two examples, machine learning model performance is also affected by how the features are represented. Feature preprocessing is a step in machine learning pipelines where we reshape the features in a manner that makes the data set easier for models to classify. For example, if we're working on a harder version of the "hill" vs. "valley" signal processing problem with noise: And we apply a "tuned" random forest to the problem: import pandas as pd import numpy as np from sklearn.ensemble import RandomForestClassifier from sklearn.decomposition import PCA from sklearn.pipeline import make_pipeline from sklearn.cross_validation import cross_val_score hill_valley_noisy_data = pd.read_csv('https://raw.githubusercontent.com/rhiever/Data-Analysis-and-Machine-Learning-Projects/master/tpot-demo/Hill_Valley_with_noise.csv.gz', sep='\t', compression='gzip') cv_scores = cross_val_score(RandomForestClassifier(n_estimators=100, n_jobs=-1), X=hill_valley_noisy_data.drop('class', axis=1).values, y=hill_valley_noisy_data.loc[:, 'class'].values, cv=10) print(cv_scores) [ 0.52459016 0.51639344 0.57377049 0.6147541 0.6557377 0.56557377 0.575 0.575 0.60833333 0.575 ] print(np.mean(cv_scores)) 0.578415300546 We'll again find that the "tuned" random forest averages a disappointing 57.8% cross-validation accuracy. However, if we preprocess the features---denoising them via Principal Component Analysis (PCA), for example: cv_scores = cross_val_score(make_pipeline(PCA(n_components=10), RandomForestClassifier(n_estimators=100, n_jobs=-1)), X=hill_valley_noisy_data.drop('class', axis=1).values, y=hill_valley_noisy_data.loc[:, 'class'].values, cv=10) print(cv_scores) [ 0.96721311 0.98360656 0.8852459 0.96721311 0.95081967 0.93442623 0.91666667 0.89166667 0.94166667 0.95833333] print(np.mean(cv_scores)) 0.93968579235 We'll find that the random forest now achieves an average of 94% cross-validation accuracy by applying a simple feature preprocessing step. Always explore numerous feature representations for your data. Machines learn differently from humans, and a feature representation that makes sense to us may not make sense to the machine. Automating data science with TPOT To summarize what we've learned so far about effective machine learning system design, we should: Always tune the hyperparameters for our models Always try out many different models Always explore numerous feature representations for our data We must also consider the following: There are thousands of possible hyperparameter configurations for every model There are dozens of popular machine learning models There are dozens of popular feature preprocessing methods This is why it can be so tedious to design effective machine learning systems. This is also why my collaborators and I created TPOT, an open source Python tool that intelligently automates the entire process. If your data set is compatible with scikit-learn, then TPOT will automatically optimize a series of feature preprocessors and models that maximize the cross-validation accuracy on the data set. For example, if we want TPOT to solve the noisy "hill" vs. "valley" classification problem: (Before running the code below, make sure to install TPOT first.) import pandas as pd from sklearn.cross_validation import train_test_split from tpot import TPOTClassifier hill_valley_noisy_data = pd.read_csv('https://raw.githubusercontent.com/rhiever/Data-Analysis-and-Machine-Learning-Projects/master/tpot-demo/Hill_Valley_with_noise.csv.gz', sep='\t', compression='gzip') X = hill_valley_noisy_data.drop('class', axis=1).values y = hill_valley_noisy_data.loc[:, 'class'].values X_train, X_test, y_train, y_test = train_test_split(X, y, train_size=0.75, test_size=0.25) my_tpot = TPOTClassifier(generations=10) my_tpot.fit(X_train, y_train) print(my_tpot.score(X_test, y_test)) 0.970100392842 Depending on the machine you're running it on, 10 TPOT generations should take about 5 minutes to complete. During this time, you're free to browse Hacker News, refill your cup of coffee, or admire the beautiful weather outside. In the meantime, TPOT will handle all of the work for you. After 5 minutes of optimization, TPOT will discover a pipeline that achieves 96% cross-validation accuracy on the noisy "hill" vs. "valley" problem---better than the hand-designed pipeline we created above! If we want to see what pipeline TPOT created, TPOT can export the corresponding scikit-learn code for us with the export() command: my_tpot.export('exported_pipeline.py') which will look something like: import numpy as np from sklearn.cross_validation import train_test_split from sklearn.ensemble import GradientBoostingClassifier from sklearn.pipeline import make_pipeline from sklearn.preprocessing import Normalizer # NOTE: Make sure that the class is labeled 'class' in the data file tpot_data = np.recfromcsv('PATH/TO/DATA/FILE', sep='COLUMN_SEPARATOR', dtype=np.float64) features = np.delete(tpot_data.view(np.float64).reshape(tpot_data.size, -1), tpot_data.dtype.names.index('class'), axis=1) training_features, testing_features, training_classes, testing_classes = \ train_test_split(features, tpot_data['class'], random_state=42) exported_pipeline = make_pipeline( Normalizer(norm="l2"), GradientBoostingClassifier(learning_rate=0.97, max_features=0.97, n_estimators=500) ) exported_pipeline.fit(training_features, training_classes) results = exported_pipeline.predict(testing_features) and shows us that a tuned gradient tree boosting classifier is probably the best model for this problem once the data has been normalized. We've designed TPOT to be an end-to-end automated machine learning system, which can act as a drop-in replacement for any scikit-learn model that you're currently using in your workflow. If TPOT sounds like the tool you've been looking for, here's a few links that you may find useful: TPOT repository on GitHub TPOT documentation TPOT installation instructions TPOT example code And as always, please feel free to get in touch. You can find all of the code used in this article on my GitHub. Enjoy! --- ## Why did so many Japanese families avoid having children in 1966? URL: https://www.randalolson.com/2016/05/08/why-did-so-many-japanese-families-avoid-having-children-in-1966/ Published: 2016-05-08 Categories: data visualization Tags: fertility rates, Japan, superstition Randy Olson delves into supervision and astrology to explain why so many Japanese families avoid having children in 1966. Last week, I was presenting at a conference and discussing the merits of animated visualizations vs. small multiples. On one of my slides, I presented the following chart that shows the total fertility rate (i.e., the average number of children born per woman) for the U.S.A. and Japan over a 60-year time period. After the talk, one of the audience members came up to me and asked why there was that weird blip in Japan's fertility rate in 1966. It turns out that there's a fascinating explanation -- an explanation that finds its roots in astrology and superstition. Astrology and superstition If you were born in the U.S. (or many other Western countries), you were probably assigned an astrological sign based on the day you were born: Aries if you were born between March 20 and April 19 (roughly), Taurus if you were born between April 19 and May 20, and so on. Each of these signs are associated with personality traits and various other features. The Japanese use a similar astrological system, but one instead based on the Chinese zodiac. Along with assigning an astrological beast based on your birth year, each year also has one of the Five Elements associated with it---all that dramatically affect what your astrological sign entails. Astrologers would like us to believe that our personality---and even our entire lives---are guided by these signs, but most people don't take these predictions too seriously. In 1966, however, many Japanese families were still quite superstitious---and that's why we see that blip in fertility rates in 1966. You see, 1966 was the year of 丙午 (Hinoe-Uma), or the "Fire Horse." As one source describes: Girls born in [1966] became known as 'Fire Horse Women' and are reputed to be dangerous, headstrong and generally bad luck for any husband. In 1966, a baby's sex couldn't be reliably detected before birth; hence there was a large increase of induced abortions and a sharp decrease in the birth rate in 1966. Instead of taking the risk of raising a "Fire Horse Woman," whose headstrong nature would bring bad luck for her future husband, many Japanese families avoided having children entirely in 1966. In essence, superstition was embedded so deeply in Japanese culture that we could measure its effects on a macro-population scale. Time will tell if superstition will strike again 10 years from now in 2026, the next year of the "Fire Horse" in its 60-year cycle. Given that Japanese is already below the replacement fertility rate (i.e., roughly an average of 2 children per woman), the result could be disastrous. --- ## Spurious Extrapolations: Novel and unique research abstracts URL: https://www.randalolson.com/2016/04/04/spurious-extrapolations-novel-and-unique-research-abstracts/ Published: 2016-04-04 Categories: analysis, data visualization Tags: publishing, research, spurious extrapolation, text analysis Randy Olson explores what would happen if the trend of using "novel" and "unique" in research abstracts continues upwards. Last Christmas, BMJ published a funny article exploring the mentions of positive and negative words in research abstracts over the past 40 years. I've recreated their research for two of the phrases below -- "novel" and "unique." Your eyes aren't fooling you: Over the past 40 years, researchers have started using the word "novel" so much that it appears in roughly 8% of all published research abstracts on PubMed. "Unique" has similarly grown in use -- now used in about 3% of all research abstracts on PubMed -- albeit not quite as dramatically. If you're interested in the details on how they looked up these phrases, read their supplemental information document. In the spirit of the Spurious Extrapolations series, I had to ask: If "novel" and "unique" kept growing in use at the same rate they have been, how long would it take until they were used in every research abstract? Disclaimer: Extrapolating beyond the bounds of a data set is extremely precarious, and most likely wrong. Most statisticians would recommend against extrapolating beyond a few time points outside of a data set. To compute these extrapolations, I fit polynomial regressions (Usage_Pct ~ Year + Year2) to the time series and used those models to predict when the terms would reach 100% usage. Shown in the chart above, "novel" will reach 100% usage by 2130 and "unique" by 2674. It's comforting to know that by 2674, our research will be novel and unique, just like all other research. Really, these charts are just an exaggeration of an already-silly phenomenon: Funding agencies and journals have placed considerable pressure on academics to perform "novel" research, which has in turn fooled many academics into thinking that some research paths are worthwhile simply because they're novel. Let's not forget that many of the greatest breakthroughs were achieved not because they were "novel," but because they built on the findings of hundreds of scientists from the past. That's how science works, and that's the kind of science we should encourage. Oh, and seriously: Please stop using the word "novel" when describing your research. We know. It's research. --- ## Spurious Extrapolations: What if U.S. college tuition costs keep rising? URL: https://www.randalolson.com/2016/03/26/spurious-extrapolations-what-if-u-s-college-tuition-costs-keep-rising/ Published: 2016-03-26 Categories: analysis, data visualization Tags: college tuition, spurious extrapolation, statistics Randy Olson embarks on a nonsensical journey exploring what would happen if U.S. college tuition costs keep rising at the rate they have been. For this post, I'm going to test run a new post series called Spurious Extrapolations, where I extrapolate time series far beyond reason and envision what would happen if the trend continued. Let me know what you think of the series in the comments. A couple years ago, I wrote an article showing that modern college students in the U.S. have to work twice as hard as their parents to pay off their college tuition. As luck would have it, the NCES recently updated their database to 2012, so I thought I'd revisit the study to see how college tuition costs compared to minimum wage up through 2012. Sure enough, the same trend holds: The average college student in 2012 has to work 2.5x as many hours on minimum wage as the average college student in 1987 to pay for the same education. It's no wonder our parents think that working your way through college is so easy: It was easy back in the day in comparison! Nowadays, soaring college tuition costs combined with stagnating minimum wages have helped cause the worst student debt crisis the U.S. has ever faced. But this is all old news, right? That's when I got to thinking: What if the U.S. minimum wage and college tuition costs keep rising at the same rate they have been? If you've followed my blog for any length of time, then you know I wasn't content just to ponder. I fit a linear regression to the data points above, then used the regression parameters to extrapolate the trend beyond 2012. Disclaimer: Extrapolating beyond the bounds of a data set is extremely precarious, and most likely wrong. Most statisticians would recommend against extrapolating beyond a few time points outside of a data set. Now, let's dispense with those tiresome warnings and have some fun. Let's extrapolate the trend of rising tuition costs to 2100: As the college students of 2100 ring in the new century, the average student will have to work over 3,100 hours on minimum wage to pay off one year of their college tuition. Let's put that into perspective: If they spread those 3,100 hours evenly over the year, they would have to work 60 hours every week just to pay for tuition! Forget Netflix and chill: how about nap time and pay off my college bill? But wait, it gets worse. So much worse. What if we focus on one of the most expensive public universities in the U.S. -- say, Pennsylvania State University, where the Pennsylvania legislature refuses to raise the minimum wage beyond Federal mandates? Here's the trend for PSU between 1987 and 2012: PSU students were already in dire straits in 1987. In 2012, PSU students had to work over twice as many hours on minimum wage as the average modern college student to pay off one year of their college tuition. Now, what if we extrapolate PSU's college tuition costs to 2100? Holy smokes! PSU students in 2100 are stuck working over 7,000 hours on minimum wage to pay off just one year of tuition. Again, let's put that into perspective: There are 24 hours in a day, 7 days in a week, and 52 weeks in a year, so that means we have about 8,700 hours to work with each year. PSU students of the future would have to spend 4/5 of their year working on minimum wage to pay off that year's college tuition. In this dystopian future, PSU students would wake up at the crack of noon to go to classes for a few hours. After that, they'll rush off to their 19-hour-a-day work shift, then maybe try to put in a couple hours of sleep before it all starts over again. Maybe they'll get lucky and sneak in a few minutes to eat, but that's unlikely because their entire paycheck is going toward the tuition bills. Perhaps we're destined to repeat history, if we live long enough. Obviously, these extrapolations were all made in jest, and we shouldn't really believe that college tuition costs are going to spiral out of control that badly. But maybe, just maybe... it'll make us think twice about the underlying problems that created that dangerous upward trend in the first place. --- ## The correct way to use pie charts URL: https://www.randalolson.com/2016/03/24/the-correct-way-to-use-pie-charts/ Published: 2016-03-24 Categories: data visualization, tutorial Tags: data visualization, pie charts Randy Olson defends the infamous pie chart and explains the correct way to use them. Pie charts are the most widely berated chart in data visualization. Many articles have been written over the years describing why pie charts are bad, and why we should no longer use them. Even key members of the data visualization community consider using a pie chart equivalent to using incorrect grammar. In short, pie charts are the comic sans of data visualization: Many would agree that if you use a pie chart, you immediately betray your lack of data visualization knowledge. I believe that the reason so many people dislike pie charts is because they're so frequently misused by amateur practitioners. Furthermore, I believe that pie charts still play an important role in data visualization, and I'm going to make my case in favor of pie charts below. Along the way, I will explain the correct way to use pie charts; please take note and share this information with your colleagues so we can salvage the pie chart's reputation. The advantages of pie charts From my point of view, pie charts have two major advantages over their alternatives: Pie charts are easy to understand. Even readers who have never taken a statistics course can look at a pie chart and immediately understand what it's trying to show. This is a vital factor if you are making data visualizations for public consumption. Pie charts easily communicate a simple proportion. If all you need to communicate is that one category (or the sum of a few categories) represents a simple proportion of a whole, then pie charts will excel at this task. I will demonstrate these points with a few examples below. First, let's cover the three most common mistakes that designers make when using pie charts. Make sure your parts sum to a meaningful whole By far, the most common mistake with pie charts is representing parts that don't sum to a meaningful whole. "What is a meaningful whole?," you ask. Let's make this concrete with an example. Say we're putting together a presentation for our boss and want to demonstrate the popularity of three programming languages as indicated by their Google search frequency. We make the pie chart below and move on with our slides. What's wrong with this pie chart? That's right: R, JavaScript, and Python don't represent every programming language out there, yet by making them the only "pieces of the pie," that's exactly what we're claiming! In other words, our parts don't sum to a meaningful whole that represents all Google searches for all programming languages. To fix this problem, we have to add an "Other programming languages" part to the pie chart, as we have below. Now we're properly visualizing the relative popularity of our three programming languages. With the pie chart above, we can easily make the point that R, JavaScript, and Python together represent a little less than 1/4 of all Google searches for programming languages. Note that we're using the pie chart to make statements about simple proportions of the whole: 1/4, 1/3, 1/2, etc. are fine as simple proportions, but don't use pie charts to communicate a specific percentage, for example, 32.3333%). Collapse categories down to 3 or fewer categories that matter Now let's say we want to provide a broader picture of all the programming languages that people search for on Google. We make the pie chart below to show the percentage breakdown for all of the programming languages that we've been tracking. What's wrong with this pie chart? Right again! We went overboard and showed way too many categories at once. Pie charts are not designed to communicate multiple proportions, so we should collapse our categories down to only a handful that really matter -- 3 or fewer is the general rule of thumb. This is where we really have to think about the message that we want to communicate with this pie chart. Do we really want to show all of the programming languages, or do we want to only focus on a few of them? In this case, we decide that we really only care about Java, PHP, and Python, and collapse the other categories into an "Other" column. Now we have a very clear message with our pie chart: Java, PHP, and Python together represent nearly 1/2 of all Google searches for programming languages, and Java represents roughly 1/4 of all searches alone. Note that we shouldn't try to use the pie chart to compare between Java, PHP, or Python: Pie slices are notoriously difficult to compare directly, especially if we ask our uncle who always takes the bigger slice of pumpkin pie during Thanksgiving dinner. If we want to compare the proportions, then we should use a bar chart instead. Always start your pie charts at the top One final note: Our readers will typically start reading our pie charts from the top of the circle -- the 0‎° mark. We should never violate our readers' expectations by starting the parts at any other section of the circle, even if it makes the pie chart look like Pac-Man. Fortunately, most data visualization software starts pie charts at the 0‎° mark. But in case you find software that doesn't: you've been warned! Recap Now that we've walked through a few visualization exercises, let's recap what we've learned about pie charts. The parts must sum to a meaningful whole. Always ask yourself what the parts add up to, and if that makes sense for what you're trying to convey. Collapse your categories down to three or fewer. Pie charts cannot communicate multiple proportions, so stick to their strengths and keep your pie charts simple. Always start your pie charts at the top. We naturally start reading pie charts at the top (the 0‎° mark). Don't violate your reader's expectations. Pie charts are useful for representing a simple proportion of a whole, and can easily be interpreted by expert and novice alike. Keep these tips in mind the next time you need to communicate a simple proportion to a general audience. If you liked what you saw in this post and want to learn more about data visualization, come take a look at my Python data visualization video course that I made in collaboration with O'Reilly. In just one hour, I will cover these topics and much more, which will provide you with a strong starting point for your career in data visualization. --- ## Why posts get removed from /r/DataIsBeautiful URL: https://www.randalolson.com/2016/03/18/why-posts-get-removed-from-rdataisbeautiful/ Published: 2016-03-18 Categories: data visualization, reddit Tags: community management, data visualization, reddit, subreddit Randy Olson delves into the data to show why posts get removed by the moderators of /r/DataIsBeautiful. I've been a moderator of /r/DataIsBeautiful -- one of the largest online communities dedicated to data analysis and visualization -- for the past 2 1/2 years. During that time, I've reviewed thousands of data visualizations created by amateurs and professionals alike. (For those not in the know: Moderators on Reddit volunteer their time to help run the subreddits, remove spam, enforce posting rules, and various other tasks to keep the subreddits on-topic and spam-free.) Although moderating on Reddit is often a thankless job, my experience on /r/DataIsBeautiful has provided me a unique perspective on the data visualization world. For this post, I thought it would be a fun exercise to visualize how /r/DataIsBeautiful's posting rules have evolved over time. After all, it only seems appropriate that a /r/DataIsBeautiful moderator would analyze and visualize their own community, right? /r/DataIsBeautiful currently has five core posting rules (1-5), and three experimental rules (6-8): These posting rules try to provide objective criteria for what makes an appropriate /r/DataIsBeautiful post, and generally make sure that: The post actually contains a data visualization The post gives appropriate credit to the data visualization's creator The post isn't obviously and maliciously misleading Evolution of /r/DataIsBeautiful's posting rules I've always been curious about the relative importance of these rules over time, so I analyzed the Reddit comment cache on BigQuery, parsed out the official moderator comments that /r/DataIsBeautiful moderators made when removing posts, and binned them by month. I was able to analyze the comments between January 2013 and January 2016, which provides a unique perspective on the subreddit before and after it became a default subreddit in early 2014. By far, the two biggest reasons for posts being removed on Reddit are: failing to link to something that includes a data visualization, and failing to properly credit the original author of the visualization Today, those two reasons constitute roughly 50% and 40% of all post removals, respectively. Interestingly, ever since it defaulted /r/DataIsBeautiful has been receiving increasingly more posts that don't even include data visualizations. I believe this trend highlights the importance of an active moderation team as a community grows: As /r/DataIsBeautiful's subscriber numbers climbed from the hundreds of thousands into the millions, the moderators were there to help the community stay on track and share relevant content. You'll also notice that there was a stint in 2014 where the moderation team required that all posts linking directly to images must link to PNGs (due to text quality issues with JPEGs), but that rule was replaced when we enacted the rule requiring that all posts link directly to the original source. Since only Original Content creators can post direct links to images now, we decided that it was best to allow them to decide how they wanted to share their content. You may also notice the lack of the "no political posts except for Thursdays" rule, which was introduced in February 2016. Since February 2016's comment data is not yet available, I'll have to leave that analysis for a future post. Post removal counts For completeness, I've also included the raw counts for each post removal reason below. You'll likely notice the spike in "original source" removals in late 2014, which was due to the /r/DataIsBeautiful mod team more strictly enforcing directly links to the original source. The community took a couple months to get used to the rule, but eventually returned to normal. For the most part, posting rules 3-8 deal with edge cases: Posters confusing infographics with data visualizations, sensationalized post titles, or reposts of links that were already shared recently. While theserules are nonetheless important, this analysis has shown us that only two of the eight posting rules play the most important role in keeping /r/DataIsBeautiful on track. We'll be sure to continually review and revise these posting rules as the /r/DataIsBeautiful community grows. If you have ideas on how we can improve the community -- through posting rules or otherwise -- please reach out to us by modmail. --- ## What data visualization tools do /r/DataIsBeautiful OC creators use? URL: https://www.randalolson.com/2016/03/11/what-data-visualization-tools-do-rdataisbeautiful-oc-creators-use/ Published: 2016-03-11 Categories: data visualization, reddit Tags: data visualization, reddit, tools Randy Olson delves into Reddit comment data to find out what data visualization tools /r/DataIsBeautiful OC creators use. One of the most common questions that newcomers to data [science/visualization/analysis] ask is: "What tools should I use to create data visualizations?" While I always recommend learning design principles before tools, I thought I'd take a stab at answering that question by analyzing what tools the /r/DataIsBeautiful community uses. For the uninitiated, /r/DataIsBeautiful is an online community dedicated to data analysis and visualization, where people post and discuss various data visualizations from around the web. Sometimes /r/DataIsBeautiful community members create and share their own data visualizations -- called "OC," or Original Content -- which I have always found to be a great source of ideas and inspiration. As part of the /r/DataIsBeautiful posting rules, every OC contributor must include a comment on their post describing the data source(s) and tool(s) they used to create their data visualization. Thus, analyzing their tool usage over the years was a fairly simple n-gram analysis of all comments made by OC contributors on /r/DataIsBeautiful that mention the word "tool." For this article, I analyzed thousands of comments made by OC contributors to /r/DataIsBeautiful between January 2014 and January 2016. (Unfortunately, it was difficult to parse out mentions of the "R" language with the n-gram analysis, so we'll have to use ggplot2 as a proxy.) The most popular tools on /r/DataIsBeautiful are: Tool Free? Requires programming? Typical uses Excel Paid No Basic data analysis and visualization Python Free Yes General-purpose scripting language that is typically used for data scraping, cleaning, and wrangling D3.js Free Yes JavaScript-based library for interactive data visualization on the web Tableau Paid, with limited free option No Advanced interactive data visualizations for the web ggplot2 Free Yes Advanced data visualization library for the R scripting language R Free Yes Scripting language designed for statistical analysis, modeling, and data visualization matplotlib Free Yes Python-based visualization library for making basic data visualizations As expected, Excel dominates the list as the primary tool that most beginners use: In this case, there have been at least 643 OC data visualizations on /r/DataIsBeautiful that were made with Excel. Excel is a great tool to start with, but you should eventually move on to more advanced tools that allow you to programmatically generate visualizations such as matplotlib/Seaborn, D3.js, or ggplot2. If programming isn't your forte, Tableau is a much better option than Excel. Here's descriptions for the rest of the tools: Tool Free? Requires programming? Typical uses JavaScript Free Yes Scripting language for the web Highcharts Free for non-commercial projects Yes JavaScript-based library for programmatically creating interactive data visualizations for the web; easier to use but less flexibility than D3.js Datawrapper Free No Basic online interactive visualizations Gephi Free No Network visualization Plotly Free No Web-based GUI for creating interactive data visualizations CartoDB Free (limited) No Web-based tool for creating interactive online maps Seaborn Free Yes Python-based visualization library for advanced statistical data visualization Matlab Paid Yes Powerful analysis, modeling, and data visualization tool Google Charts Free Yes Simple JavaScript-based visualization library for creating interactive online visualizations Leaflet.js Free Yes Simple JavaScript-based visualization library for creating interactive online map visualizations LaTeX Free Yes Document preparation system that it somehow used to create visualizations (???) Google Fusion Free No Web-based tool for creating interactive online data and map visualizations Bokeh Free Yes Python-based visualization library for creating interactive data and map visualizations I was also curious about temporal trends in library usage, so I grouped the tool mentions by year and plotted them below. GUI-based visualization tools such as Tableau and Gephi are seeing steady growth, whereas Python and matplotlib (oddly) seem to be waning in relative popularity. D3.js and ggplot2 are similarly experiencing steady growth, although I should note that 2016's counts are only based on January 2016's comments and may change by the end of the year. We'll have to revisit these trends come 2017. Hopefully that answers all of your data visualization tool-related questions! If you have any more questions or concerns, please leave them in the comments. How to download the comments I analyzed If you'd like to repeat this analysis yourself, run the following SQL statement on the Google BigQuery database. SELECT body, created_utc FROM [fh-bigquery:reddit_comments.2016_01], [fh-bigquery:reddit_comments.2015_12], [fh-bigquery:reddit_comments.2015_11], [fh-bigquery:reddit_comments.2015_10], [fh-bigquery:reddit_comments.2015_09], [fh-bigquery:reddit_comments.2015_08], [fh-bigquery:reddit_comments.2015_07], [fh-bigquery:reddit_comments.2015_06], [fh-bigquery:reddit_comments.2015_05], [fh-bigquery:reddit_comments.2015_04], [fh-bigquery:reddit_comments.2015_03], [fh-bigquery:reddit_comments.2015_02], [fh-bigquery:reddit_comments.2015_01], [fh-bigquery:reddit_comments.2014] WHERE LOWER(BODY) LIKE "%tool%" AND subreddit == "dataisbeautiful" --- ## Major League Baseball home run leaders, 1871-2016 URL: https://www.randalolson.com/2016/03/10/major-league-baseball-home-run-leaders-1871-2016/ Published: 2016-03-10 Categories: data visualization, python, tutorial Tags: animated visualizations, data visualization, major league baseball Randy Olson highlights one example where animated data visualization outperforms static visualizations. Earlier this week, a Reddit user shared a fascinating animated data visualization showing the MLB home run leaders from the past 200+ years. I found this visualization especially interesting because it was one of the few examples where I've seen an animated data visualization effectively tell a story that a static visualization couldn't tell. I've recreated the user's visualization with Python below. Click here for a gfycat version We see the early battle for the home run throne in the late 1800's, then a period of stagnation in the early 1900s. As one Redditor explained: The reason why the records remain fairly stagnant from 1903-1920 is because there was a Dead-ball era in the MLB. During the "Dead Ball Era," ballparks had ridiculously large dimensions, balls were used until they weren't useful anymore, and many pitchers "doctored" the ball by spitting on it or covering it in tobacco. Shortly thereafter, we see why Babe Ruth was such a sensation in the 1920s, as his record skyrocketed to 714 home runs over his career. Ruth held that record until the 1970s, when Hank Aaron dethroned him and held onto the record for another 40 years. There's a lot more to discover in this animated visualization, but I'll leave that as an exercise to the reader. The downside of this visualization Of course, the most obvious criticism of this visualization is that the lines represent different people over time, which can be disorienting to some viewers. It's important to remember that this visualization is meant to show the evolution of home run records over time, and not necessarily the home run records of any particular individual. One possible way to overcome this shortcoming is to take a cue from xkcd and assign each player their own line: However, we would likely have to limit the number of players we visualize at once, and would likely only be able to show one or two dominant players during each time period. Furthermore, since career home run records only go up over time, we would quickly see the 400+ home run range filled with several player's records. Perhaps an xkcd-like version can be made for yearly home runs of dominant players, but I'll leave that as an exercise for the future. How do I remake this visualization? Below is the Python code that I used to generate the animated visualization. Once you've generated all of the individual frames, you'll have to stitch them together with a program such as ffmpeg or Camtasia. import matplotlib.pyplot as plt import pandas as pd # This is my custom matplotlib style -- feel free to reuse it plt.style.use('https://gist.githubusercontent.com/rhiever/d0a7332fe0beebfdc3d5/raw/223d70799b48131d5ce2723cd5784f39d7a3a653/tableau10.mplstyle') mlb_data = pd.read_csv('http://www.randalolson.com/assets/2016/03/mlb-home-run-leaders-static.csv', sep='\t') mlb_data.set_index('year', inplace=True) # For every year (except the first 5)... for year in mlb_data.index.unique()[5:]: # Each year gets its own figure plt.figure(figsize=(6, 9)) # Subset the data to only the data leading up to the current year subset = mlb_data.loc[mlb_data.index --- ## Revisiting the vaccine visualizations URL: https://www.randalolson.com/2016/03/04/revisiting-the-vaccine-visualizations/ Published: 2016-03-04 Categories: data visualization, python, tutorial Tags: data visualization, infectious disease, vaccines Randy Olson revisits the popularized vaccine visualizations published by WSJ and provide a fresh look on the data. Last year, the vaccination debate was all the rage again. "Pro-vaxxers" were loudly proclaiming that everyone should get vaccinated and discussing the science behind it, and "anti-vaxxers" were casting their doubts and still refusing to get vaccinated for personal reasons. Around that time, The Wall Street Journal released a brilliant series of heat maps showing infection rates for various diseases over time, broken down by state. These heat maps easily demonstrated one of the most important facts in the vaccination debate: Time and time again, vaccines work. Today, I would like to revisit the WSJ's heat maps through the lens of a data visualization practitioner. In particular, I would like to show how these heat maps can possibly be improved upon by reviewing some basic rules of data visualization, and trying out some other methods for displaying the data. Below, I'm going to walk through four major criticisms and show how addressing them can possibly improve the original work. For the curious, I've released my notebook with the Python code used to generate the new visualizations. Categorical color palettes should not be used to display continuous values Perhaps one of the most straining issues with the original WSJ heat maps was their use of a custom categorical color palette to display the infection rates. The palette runs through most of the colors of the rainbow at seemingly-random intervals. It's possible that they calculated the quantiles to determine the ranges for the color bins (as they should!), but that wasn't indicated in their methodology. In any case, it's rarely a good idea to use multiple colors to display a single continuous variable. Here, all we want to do is use color to show the infection rates for each year. If we use more than one color, our readers have to constantly refer back to the legend to figure out what each color means, which is an unnecessary cognitive strain on our reader. Instead, we should use a single-color sequential palette, where lighter shades indicate lower values and darker shades indicate higher values. I've reworked the Polio heat map to do just that below. One exception to this "rule," of course, is diverging color palettes. If there is a clear divide in our continuous variable -- for example, if we're displaying gains and losses for a company -- then it could be appropriate to use a diverging color palette with one color to represent gains (values >= $0) and another to represent losses (values Just for fun, I recreated the same chart above for Measles so we can compare it to the originals on WSJ. Multi-hue color palettes should take color blindness into account Color blindness is probably one of the most-overlooked issues in data visualization, and the WSJ heat maps are a great example. I ran the WSJ heat map above through a color blindness simulator for red-green color blindness -- the most common form of color blindness -- and below is the result. Disastrous! Much of the color gradient is lost in some yellow/grey abyss, and the dark purple colors represent low values whereas the lighter yellow and dark grey colors represent higher values. This color palette survives better than most and the main message is still (mostly) communicated, but the WSJ color palette is certainly far from ideal here. For comparison, I ran my rework from above through the same red-green color blindness simulator. As we can see, the simple sequential color palette is practically unaffected by this form of color blindness. Problem solved! The main lesson here is that we should always run our color palette through a color blindness simulator before committing to it. Roughly 5% of our audience will experience our data visualizations through that lens. Color can't display specific values very well One of the major drawbacks of heat maps is that they rely on color to communicate the specific values in each cell. While it's not always important to display a precise value, there can sometimes be important trends hiding in these small differences. For that reason, I reworked the Polio heat map into a simple line chart below, where each light line is a state and the dark line is the median value between all the states for each year. The above chart isn't too useful, and the data is too messy to make much sense of the state-by-state trends. However, the decline in infection rates after the introduction of the vaccine is abundantly clear even in this case. No post of mine is complete without small multiples, so let's give that a try. Below, each state has its own chart, and all 50 states (+ D.C.) are put on the same time axis. Each line tells its own story, and these are stories that were masked in the heat maps. Small multiples allow use to see specific state-by-state trends, for example, Polio outbreaks were already on the decline in South Dakota even before the introduction of the Polio vaccine. Meanwhile, Polio outbreaks were at their worst in New Hampshire just prior to the introduction of the Polio vaccine, which made short order of Polio immediately thereafter. We should always ask ourselves when designing data visualizations: Do we care about the broader story, or the smaller stories? In this case we could go either way, but the direction we go depends on the story we want to tell. Sometimes you can show too much data Another fair criticism of all the data visualizations shown so far is that they show too much data. After all, the main message of the WSJ heat maps was simple: When introduced to human populations, vaccines work. There's no need to show the state-by-state trends then; in fact, we may be overwhelming our reader by providing too much data that doesn't get right to the point. For example, what happened with Polio in Utah, with the infection rate more than doubling after the introduction of the Polio vaccine? Or what about South Dakota, where Polio seems to have been mostly eliminated even before the vaccines were made available? These outliers are distractions to the overall trend. We can overcome these distractions by applying a simple statistical analysis to the data, and show the overall trend with confidence bounds. Below, I've done just that by plotting the median Polio infection rate across all states (dark line) with bootstrapped 95% confidence intervals (shaded area). By summarizing the data with some basic statistics, we've removed the distractions and gotten straight to the point: Overall in the U.S., Polio outbreaks were on the rise from the 1940s onward. Right at the introduction of the Polio vaccine in 1955, we immediately saw a decline in Polio outbreaks until it was practically eliminated in the 1960s. Again, we should always consider our story when designing data visualizations. If we have one clear story that we want to communicate, we should consider reducing the amount of data we show to the point that we can effectively -- and honestly -- communicate our story. There's no point in confusing our reader with unnecessary details, unless those details contain an important caveat. An aside At face value, these charts only demonstrate correlations: When vaccinations were introduced to the population, the prevalence of infectious disease decreased shortly thereafter. I believe it's important to point out here that even though I want to focus on data visualization techniques in this post, the science behind vaccination is not up for debate, and these charts are in fact demonstrating a proven causal relationship. Please don't waste your time typing out "correlation != causation" in the comments. Conclusions To wrap up, these are the lessons we've drawn from revisiting the popularized vaccine visualizations: Use sequential color schemes when presenting continuous values Consider color blindness before committing to a color scheme When presenting specific values is important, don't use color to represent those values Only show enough data to effectively and honestly tell your story If you liked what you saw in this post and want to learn more, check out my Python data visualization video course that I made in collaboration with O'Reilly. In just one hour, I will cover these topics and much more, which will provide you with a strong starting point for your career in data visualization. --- ## Analyzing MMA: The Ultimate Fighting Championship URL: https://www.randalolson.com/2016/02/25/analyzing-mma-the-ultimate-fighting-championship/ Published: 2016-02-25 Categories: data visualization Tags: data visualization, mma, ufc Randy Olson analyzes MMA fighters and fights in UFC, the Ultimate Fighting Championship. For the past 7 years, I've been a fan of MMA, and especially the larger Ultimate Fighting Championship events that take place around the world. For the uninitiated, MMA fights pit two professional fighters against each other who often have very different fighting styles, which leads to some exciting and often unexpected conclusions. See below for an example. Earlier today, I ran across a data set containing all of the UFC fighters and fights that have occurred from the beginning of the UFC in 1993 to 2016. I couldn't help but spend a few hours looking for trends in the data set, which I've shared below. UFC fights are lasting longer UFC fights are lasting significantly longer than they used to. Back in the early days, fights would last only a few minutes on average, regularly ending in savage knockouts that attracted so many of the early UFC fans. In fact, one of the shortest fights in UFC history took place in 1996, when Don "The Predator" Frye knocked out Thomas Ramirez in 8 seconds, prematurely ending the super heavyweight's career that night. In many ways, the increased fight lengths were due to stricter regulations forced on the UFC. In the late 90's, the UFC was forced to enact safety regulations, 5-minute rounds, and weight classes, which is immediately noticeable in the data. Fight times only continued to rise from there, averaging about 15 minutes now in 2016. Far more UFC fights are decided by the judges Even though it was the knockouts that attracted so many early fans, the large majority of the early fights ended in submissions like the armbar. After the UFC reforms in the late 1990s, fights were forced to end after three 5-minute rounds (five rounds for title matches), and the winner would be chosen by a committee of judges based on their performance in the fight. Nowadays, half of all fights run the full 15 minutes and go to decision, which is really a testament to the improved safety regulations. Interestingly, despite the rule changes, KO and TKO rates haven't changed much over the years. Submissions aren't nearly as common nowadays, but fighters like Ronda Rousey are infamous for their brutally efficient armbars. UFC fighters are getting smaller In the early days of UFC, many of the fights were all about massive men duking it out in the ring. In fact, one of the earliest UFC events was titled "David vs. Goliath" because it featured several fights where heavyweight fighters were matched up against much smaller fighters. (You might be surprised how many of those fights turned out!) Once the new rules were put in place in the late 1990s, the super heavyweight class practically disappeared and the heavyweight class has been on the way out ever since. I had trouble visualizing these trends, so here's attempt #1. Below is a stacked area chart, where each color represents a UFC weight class. The darkest color represents the hulking monsters in the super heavyweight class (265+ lbs), whereas the lightest color represents the light-as-a-feather fighters in the strawweight class (115 lbs). Generally, you'll notice that the lighter classes start to fill up more of the UFC fights as time goes on, and the heavier classes all but disappear. Here's another view of the same data, this time as small multiples. Same trends, different view: The hulking beasts of the early days subsided in favor of nimbler fighters nowadays. UFC has continually had to add new, lighter weight classes to accommodate the ever-shrinking waistlines of their fighters, especially with the introduction of the female UFC fighters in 2012. Of course, these trends are reflected in the overall weights of the UFC fighters as well, with the average fighter weighing 220+ lbs (99.8kg) in the early 1990s, compared to a lean 160 lbs (72.5kg) in 2016. The average UFC fighter is quite a bit shorter nowadays too, with a drop from 6'2" (1.88m) in 1993 down to 5'10" (1.78m) in 2016. Despite the ever-shrinking fighters, UFC fights are no less exciting than they used to be. In fact, Carla Esparza -- standing at only 5'1" (1.55m) -- holds the honor as one of the best fighters in the UFC strawweight class, and it sure shows in the ring. If you'd like to explore the data yourself, head on over to the Reddit thread and give it a gander. Enjoy! --- ## Introducing TPOT, the Data Science Assistant URL: https://www.randalolson.com/2015/11/15/introducing-tpot-the-data-science-assistant/ Published: 2015-11-15 Categories: machine learning, python, research Tags: automation, data science, evolutionary computing, machine learning, pipeline, python Randy Olson introduces TPOT, an Automated Machine Learning tool intended to be your Data Science Assistant. Some of you might have been wondering what the heck I've been up to for the past few months. I haven't been posting much on my blog lately, and I haven't been working on important problems like solving Where's Waldo? and optimizing road trips around the world. (I promise: I'll get back to fun posts like that soon!) Instead, I've been working on something far geekier, and I'm excited to finally have something to show for it. Over the summer, I started a new postdoctoral research position funded by the NIH at the University of Pennsylvania Computational Genetics Lab. During my first month there, I started looking for big problems in the field of data science to take on. Science (especially computer science) is often too incremental, and if I was going to stay in academia, I wanted to tackle a big problem. It was around that time that I started thinking about the process of machine learning and how we could let machines solve problems themselves rather than needing input from humans. You see, machine learning is transforming the world as we know it. Google search engines were massively improved by machine learning, as were Gmail's spam filters. Voice assistants like Siri -- as silly as they can be -- use machine learning to translate your voice into something the computer can understand. Stock market investors make millions every day using machine learning to predict when to buy and sell. And the list goes on and on... Ever wonder how Facebook always knows who you are in your photos? They use machine learning. The problem with machine learning is that building an effective model can require a ton of human input. Humans have to figure out the right way to transform the data before feeding it to the machine learning model. Then they have to pick the right machine learning model that will learn from the data best, and then there's a whole bunch of model parameters to tweak that can make the difference between a dud and a Nostradamus-like model. Building these pipelines -- i.e., sequences of steps that turn the raw data into a predictive model -- can easily take weeks of tinkering depending on the difficulty of the problem. This is obviously a huge issue when machine learning is supposed to allow machines to learn on their own. An example machine learning pipeline, and what parts of the pipeline TPOT automates Thus, the Tree-based Pipeline Optimization Tool (TPOT) was born. TPOT is a Python tool that automatically creates and optimizes machine learning pipelines using genetic programming. Think of TPOT as your "Data Science Assistant": TPOT will automate the most tedious part of machine learning by intelligently exploring thousands of possible pipelines, then recommending the pipelines that work best for your data. An example TPOT pipeline with two copies of the data set entering the pipeline. Once TPOT is finished searching (or you get tired of waiting), it provides you with the Python code for the best pipeline it found so you can tinker with the pipeline from there. As an added bonus, TPOT is built on top of scikit-learn, so all of the code it generates should look familiar... if you're familiar with scikit-learn, anyway. TPOT is still under active development and in its early stages, but it's worked very well on the classification problems I've applied it to so far. Check out the TPOT GitHub repository to see the latest goings on. I'll be working on TPOT and pushing the boundaries of machine learning pipeline optimization for the majority of my postdoc. An example using TPOT I wanted to make TPOT versatile, so it can be used on the command line or via Python scripts. You can look up the detailed usage instructions on the GitHub repository if you're interested. For this post, I've provided a basic example of how you can use TPOT to build a pipeline that classifies hand-written digits in the classic MNIST data set. from tpot import TPOTClassifier from sklearn.datasets import load_digits from sklearn.cross_validation import train_test_split digits = load_digits() X_train, X_test, y_train, y_test = train_test_split(digits.data, digits.target, train_size=0.75, test_size=0.25) tpot = TPOTClassifier(generations=5, verbosity=2) tpot.fit(X_train, y_train) print(tpot.score(X_test, y_test)) After 10 or so minutes, TPOT will discover a pipeline that achieves roughly 98% accuracy. In this case, TPOT will probably discover that a random forest classifier and k-nearest-neighbor classifier does very well on MNIST with only a little bit of tuning. If you give TPOT even more time by setting the "generations" parameter to a higher number, it may find even better pipelines. "TPOT sounds cool! How can I get involved?" TPOT is an open source project, and I'm happy to have you join our efforts to build the best tool possible. If you want to contribute some code, check the existing issues for bugs or enhancements to work on. If you have an idea for an extension to TPOT, please file a new issue so we can discuss it. tl;dr in image format Justin Kiggins had a great summary of TPOT when I first tweeted about it: @randal_olson pic.twitter.com/ds5iTbA2oF — Justin Kiggins (@neuromusic) November 13, 2015 Anyway, that's what I've been up to lately. I'm looking forward to presenting TPOT at several research conferences in the coming months, and I'd really like to see what the machine learning community thinks about pipeline automation. In the meantime, give TPOT a try and let me know what you think. --- ## Visualizing Indego bike share usage patterns in Philadelphia (Part 2) URL: https://www.randalolson.com/2015/09/05/visualizing-indego-bike-share-usage-patterns-in-philadelphia-part-2/ Published: 2015-09-05 Categories: data visualization Tags: bike share, Indego, Philadelphia Randy Olson revisits the Indego bike share station usage data to find better ways to visualize the daily usage patterns. A couple months ago, I made an initial foray into understanding the usage patterns of Indego, Philadelphia's new bike share system. This month, I thought it'd be a fun exercise to revisit that data set to see if I could visualize the usage patterns a little better. From my previous attempt, we already know that there are clear cyclical usage patterns for many of the Indego stations: Stations in residential areas tend to empty out in the mornings and refill in the evenings, whereas stations in the business sectors do the opposite. In addition to that information, I also wanted to communicate the health of the bike stations, i.e., how often the stations are full or empty of bikes. Below is my attempt at combining both pieces of information into a single graph. Use this visualization to plan your commutes and avoid the busy hours for your local bike stations. Each sub-graph corresponds to an active Indego station, showing the average performance from morning (left) through the evening (right). The y-axis represents the percentage of time the station is full, empty, or somewhere in between during the corresponding hour. Darker reds represent fuller stations, whereas darker blues represent emptier stations. To give a specific example, let's zoom in on the 21st & Catharine station. The 21st & Catharine station is clearly a commuter station: The dominant red shades in the mornings and evenings show that it tends to fill up with bikes during non-work hours, while the creeping blue shades during the workday show the station emptying out as people ride bikes to work. In contrast, the station at 17th & Girard sees much less action: The station sits more-empty-than-full most of the time, with only a minor outflux of bikes during work hours. The nice part about these visualizations is that they show how reliably available each bike station is throughout the day: Stations like 17th & Girard are practically always ready to receive or provide a bike at any point in the day, whereas stations like 21st & Catharine are regularly bouncing between full and empty on the weekdays. So, what do you think? Does this visualization method show the station usage trends better than before? How do you think it can be improved from here? --- ## Small multiples vs. animated GIFs for showing changes in fertility rates over time URL: https://www.randalolson.com/2015/08/23/small-multiples-vs-animated-gifs-for-showing-changes-in-fertility-rates-over-time/ Published: 2015-08-23 Categories: data visualization, tutorial Tags: fertility rates, GIFs, small multiples Randy Olson demonstrates how small multiples can be used to communicate data more effectively than animated GIFs. A couple weeks ago, Stephen Holzman shared an animated GIF on /r/DataIsBeautiful that caught my eye. The GIF showed the evolution of fertility rates of the U.S. and Japan between 1947 and 2010, which starts right in the middle of the post-WWII Baby Boom and follows the gradual decline of Japan's fertility rates, which has led to somewhat of a population crisis for Japan. Although Stephen's GIF is fun to watch -- especially because the animation gives the appearance of waves rising and falling -- I couldn't help but be frustrated by the limitations of GIFs in data visualization. If we wanted to compare the fertility rates of 1980 and 2010, for example, we'd have to keep a mental snapshot of what the 1980 frame looked like for when the 2010 frame came around. Thus, comparisons of time points beyond a couple years are impossible with animated GIFs unless the viewer has photographic memory. This drawback is the exact reason that small multiples were introduced to data visualization: If we're comparing the same data in the same format between several different [times|treatments|countries|etc.], then we can visualize the data on the same scale and axes to make them easily comparable. I've long been a proponent of small multiples over GIFs, so I took Stephen's data (which is actually from the Human Fertility Database) and reworked it into small multiples. You can click on the image for a super-high-res version. Each year gets its own plot -- running from left to right -- with both country's fertility rates plotted. The total fertility rate for each year is annotated onto its corresponding plot, and color-coded according to the country. I plotted the x-axis tick labels to show the reader the age range of the plots, but only on the top and bottom rows to avoid too much repetition. Similarly, the y-axis tick labels only appear on the plots on the left. Of course, the drawback of small multiples is that we no longer see the data in the same detail as we did with the larger plots. Out of necessity, each plot in a small multiples chart must be small, simple, and have few axis ticks, which can make small multiples a poor choice if you're making a comparison where there has been little change. We can compensate for this by subsetting the data. After all, 64 years is quite a lot of data to show in one graph. What if we just looked at every five years? Now it's straightforward to compare across and within decades: 1947, 1955, 1965, etc. can easily be compared by looking down the column. By the same token, 1947 and 1950 can easily be compared by looking down the row. We still get about the same level of detail as the GIF, and maintain the overall trend of declining fertility rates in both countries as time goes on. From this chart, two major trends that are readily apparent from the data: 1) Both the USA and Japan have experienced declining birth rates since the 1940s -- Japan moreso. 2) In the past 20 years, Japanese couples have started having children later in life (after their 30s) -- so much so that in 2010, half the children born were born to parents older than 30. Which begs the question: Why show all this data if we only have two points to make? Simplifying the charts even more If the above two trends are all we wanted to show with the data, then we can simplify the charts even more by calculating summary statistics and plotting those instead. These charts take away the opportunity for the reader to glean any additional insights from the data. However, if we wanted to tell a straightforward story with charts, these would be the best ones to use. Conclusions Animated GIFs, while flashy, often make it more difficult to gain insight from data. Static charts, such as small multiples, can simplify animated GIFs to make trends in the data more apparent. Sometimes it's better to calculate summary statistics and plot those instead, especially if showing all of the data does not lend additional insight. If you liked what you saw in this post and want to learn more, check out my Python data visualization video course that I made in collaboration with O'Reilly. In just one hour, I will cover these topics and much more, which will provide you with a strong starting point for your career in data visualization. Code for the small multiples visualization I can't share the data that I used to create this visualization -- you'll have to download it from the Human Fertility Database -- but I've provided the Python code I used to generate the small multiples visualization below for education purposes. import matplotlib.pyplot as plt import pandas as pd import numpy as np plt.style.use('https://gist.githubusercontent.com/rhiever/d0a7332fe0beebfdc3d5/raw/' '223d70799b48131d5ce2723cd5784f39d7a3a653/tableau10.mplstyle') # "japan_fertility" and "usa_fertility" are pandas DataFrames with # the fertility data from the Human Fertility Database plt.figure(figsize=(12, 16)) for plot_num, ((index_japan, japan), (index_usa, usa)) in enumerate( zip(japan_fertility.groupby('Year'), usa_fertility.groupby('Year'))): ax = plt.subplot(8, 8, plot_num + 1) plt.fill_between(usa.Age.values, usa.ASFR.values, color='#1f77b4', alpha=0.7) plt.fill_between(japan.Age.values, japan.ASFR.values, color='#d62728', alpha=0.7) plt.xlim(9, 51) plt.ylim(0, 0.3) if index_japan = 2003: plt.xticks(range(10, 51, 20), fontsize=10) else: plt.xticks(range(10, 51, 20), ['']) if plot_num % 8 == 0: plt.yticks(np.arange(0.1, 0.31, 0.1), fontsize=10) else: plt.yticks(np.arange(0.1, 0.31, 0.1), ['']) plt.text(40, 0.26, usa.ASFR.sum().round(2), fontsize=10, ha='center', color='#1f77b4') plt.text(40, 0.225, japan.ASFR.sum().round(2), fontsize=10, ha='center', color='#d62728') plt.title(index_japan, fontsize=10) plt.tight_layout() plt.savefig('usa-vs-japan-fertility-rates-small-multiple.pdf', dpi=300) Note that I had to add the plot axis labels, the plot title, and a couple annotations manually. --- ## U.S. college majors: Median yearly earnings vs. gender ratio URL: https://www.randalolson.com/2015/08/16/u-s-college-majors-median-yearly-earnings-vs-gender-ratio/ Published: 2015-08-16 Categories: analysis, data visualization Tags: college major, gender, pay gap Randy Olson explores a curious correlation showing that female-dominated college majors earn less than male-dominated majors. Last year, I looked at the gender ratios across college majors and discovered an interesting-yet-spurious correlation: College majors with higher male:female ratios (i.e., with more men than women) tend to have students with higher estimated IQs. After much debate, the correlation seemed to be explained by the fact that male-dominated majors tend to be more quantitative in nature, and the IQ estimation procedure relied heavily on students' quantitative SAT score. Thus, we were only observing a gender preference for quantitative vs. non-quantitative majors. Recently, I ran across an interesting FiveThirtyEight article that analyzed the estimated median earnings for recent college graduates broken down by major. The data that they used from the American Community Survey was publicly available on their GitHub, so I decided to take a closer look at the correlation between estimated median earnings and the gender ratio for each majora. Below, each square is a college majorb scaled by the number of recent college graduatesc. The squares are colored by the general category of each major, and I annotated a handful of the most popular majors. If you'd like to look up specific majors, use FiveThirtyEight's lookup tables. The trend that's immediately apparent from this chart is that female-dominated majors make less on average than male-dominated majors. Some interesting exceptions to the trend are Nursing (90% women; $48k median earnings) and Transportation Science (12% women; $35k median earnings), where Nursing especially stands out as a relatively lucrative major despite being primarily women. Despite the outliers, we're still left to wonder: Is it really true that women earn less than men -- even for college graduates? I decided to investigate this issue a little further. Of course, the first step is to make sure that our eyes aren't fooling us, so I fit a linear regression onto the data and weighted the majors by their number of recent graduates. Sure enough, there's a clear and significant negative correlation between a college major's median year earnings and gender ratio. Let's try to find out why. Why do female-dominated majors earn less? If you look through the list of female-dominated majors, you'll notice that many of them can be quite difficult to find a job with because they're so competitive: There are far more graduates than jobs in many of the fields. Could female-dominated majors be earning less because they couldn't find a job after college? The unemployment rate for each major was included in the data set, so let's take a look at that correlation below. For the most part, this idea seems to be busted: There is no significant correlation between the gender ratios and unemployment rates of college majors, and most majors are sitting around the national unemployment rate of 5.5%. What about underemployment? In his FiveThirtyEight article, Ben Casselman wrote that drama majors (which tend to be female more often than male) are more likely to end up waiting tables than using their degree. Underemployment rate was also available for each major, so let's take a look. Yet again, our idea has been busted: There is only a weak correlation between the gender ratios and underemployment rates of college majors, so the typical narrative of useless college degrees doesn't seem to explain why male-dominated majors earn more than female-dominated ones. I was a bit lost at this point until I remembered my controversial post from last year. Back to an old study... If you read my blog last year (or read the introduction to this post!), you may remember the correlation between quantitative SAT scores (and IQs) and gender ratios of college majors. What if we matched up the quantitative SAT scores -- which can be used as a proxy for how quantitatively-focused a major is -- with the median earnings for each majord? (Again, the majors are scaled by the number of recent graduates to provide a sense of relative popularity.) (For the stats nerds: R^2 = 0.624) Perhaps this all makes sense now: It seems possible that male-dominated majors -- such as Engineering, Physics, and Computer Science -- earn more than female-dominated majors because male-dominated majors are often more quantitative in nature. These quantitative majors are often employed by large companies to design products, perform data analysis, manage the company, etc., and their salaries are higher to match the responsibilities of the job. It's another question of whether businesses and governments should value the services provided by quantitative majors more than, say, Education majors, but I'll leave that discussion for another day. Disclaimer Of course, correlation != causation, and I'm not trying to make any causative inferences here. We would need to perform controlled experiments if we wanted to establish a causal relationship. I'm simply building a plausible narrative around these correlations in case anyone wanted to explore these correlations in more depth in the future. Summary Female-dominated majors tend to earn less than male-dominated majors This correlation isn't explained by the employability of the majors It seems plausible that male-dominated majors are usually paid more because they are more quantitative in nature, which large companies tend to value highly Alternative explanations So, what do you think? Do you think this explanation holds, or do you think there's another explanation at work here? I'll do my best to describe plausible alternative explanations in this section. One of the commenters mentioned that we should view the correlation in the opposite direction: "Men are more driven by income when choosing majors than women. Women value attributes other than income more than men." Technical notes a There are plenty of caveats that come with this data set, and I advise reading up about them in the FiveThirtyEight article. b I excluded Petroleum Engineering from this analysis for clarity, which is a small outlier major (only 2,339 recent graduates) with a median salary of $110k and only 12% women. c "Recent" graduates are graduates under age 28, which is roughly within five years of graduation on a normal schedule. d I performed this merger manually and did my best to match the majors when their name did not match exactly. Several majors were dropped since quantitative SAT scores were only available for a subset of majors. You can find that data set here. --- ## Analyzing the health of Philadelphia's bike share system URL: https://www.randalolson.com/2015/08/15/analyzing-the-health-of-philadelphias-bike-share-system/ Published: 2015-08-15 Categories: analysis, data visualization Tags: bike share, Indego, machine learning, modeling, Philadelphia Randy Olson analyzes Philadelphia's bike share system to explore how often the bike stations are available for use. Last month, I wrote about my initial attempts to model and predict the usage patterns of Indego, Philadelphia's new bike share system. To recap: If you've ever used a bike share before, you know that one of the biggest fears is coming up to an empty bike share station when you need a bike. (Or similarly, coming up to a full station when you need to drop a bike off.) To help abate those fears, I've started monitoring the Indego bike usage API to see if I could model and predict when the bike share stations are most likely going to be empty or full. A couple weeks after my post, another team of data analysts made an interesting suggestion to the bike sharers in Washington, D.C.: If you run into an empty or full bike share station, just wait 5 minutes and there should be an opening for you. Since I had the Indego data on handa, I decided to take a look at Philadelphia's bike share to see if the same advice holds here. Below, I analyzed each Indego station individually and plotted the percentage of the time each station recovered from becoming empty or full within 5 minutes. Each dot represents a station; if you hover over one of the dots, the plot will tell you the name of the station, the percents, and the number of times each station became empty or full since July '15. The takeaway from this plot is that if you see an empty or full Indego bike station, you're better off finding another station than waiting 5 minutes. In fact, in the past couple months you typically had to wait about 30 minutes at a station before it was no longer full or empty -- with shorter wait times during peak hoursb, of course. Thankfully, most of the Indego bike stations are quite healthy. Below, I plotted the breakdown of how often each station is healthy (not empty nor full), empty, or full and sorted the stations by how the percentage of time they're healthy. The majority of the Indego bike stations have an uptime of 90% or higher, with the station at 18th & Fernon topping the list with a near-100% uptime. By far, the biggest problem Indego bike stations face is running out of bikes during the peak hours. The Children's Hospital of Philadelphia (CHOP) station, for example, frequently runs out of bikes in the evening when the hospital workers are heading back home. Similarly, commuting stations in the residential neighborhoods of Philadelphia -- such as 18th & Washington Ave -- have a strong tendency to fill up shortly after work hours. To provide a better perspective on the Indego bike station downtimes, I plotted each station below. Again, each dot represents a station and you can hover over the dot to see the station names and statistics. The better-performing stations are on the bottom left of the chart, with near-zero downtimes. Again, the critical issue of running out of bikes is highlighted in this chart: Some stations, such as 8th & Market and 18th & JFK, spend upwards of 25% of their time without a single bike. It's fair to assume that most of this downtime is during the evenings and overnight, but it's an alarming statistic nonetheless. Despite these difficulties, I think the above findings are a great sign for the future of our burgeoning bike share program. There's currently more demand for the bike share than it can sometimes provide, which means there's tremendous opportunity for bike share growth in Philadelphia. In the meantime, if you'd like to get the upper hand on the other bike sharers and find out the best times to take your bike routes, feel free to analyze my cache of the Indego usage data below. Technical notes a I have made my cache of the Indego bike share data set available online here. (Note: Sorry, I had to take down the cache, but I've left this text here for historical records.) The data is in tab-separated format, compressed with gzip, and contains all of the data from the API, captured every 5 minutes. The cache auto-updates every 5 minutes after the hour, every hour. Please don't send requests for the cache more than one time per hour; I will have to take it down if my web site slows down due to high volumes of requests. b The peak usage hours for the Indego bike share are typically 8-10 AM and 4-6 PM ET on the weekdays. You can read more about the Indego usage patterns in my previous post. --- ## Can the Name Age Calculator guess how old you are? URL: https://www.randalolson.com/2015/08/13/can-the-name-age-calculator-guess-how-old-you-are/ Published: 2015-08-13 Categories: analysis, data visualization Tags: age calculator, baby names, FiveThirtyEight, Social Security Administration Randy Olson demonstrates how you can guess someone's age when all you know is their name. Can you guess someone's age when all you know is their first name? That was the crazy idea behind one of FiveThirtyEight's articles last year, and their surprising answer is, "Yes." The idea behind guessing someone's age based on their name is simple: There exists an openly-accessible database of when everyone in the U.S. was born and what they were named. If we add in the U.S. actuarial tables, which provides us with estimates of how many men and women born each year are still alive today, we can get a rough estimate of the age distribution of people with a given name. Take Joseph, for example: As the FiveThirtyEight team writes: Joseph has been one of the most enduring American names; it’s never gone out of fashion. So knowing that a man is named Joseph doesn’t tell you very much about his age. The median living Joseph is 37 years old, and the interquartile range (that is, the range spanning the 25th through 75th percentiles) runs from 21 to 56. In other words, a quarter of living Josephs are older than 56 and a quarter are younger than 21; the rest are somewhere in between. Not very helpful. Not very helpful indeed. However, this method shines for most names because names tend to go through short trend cycles. Brittany, for example, was popular name for girls in the 1980s and 1990s, but waned in the 2000s. Since most of those Brittanys are still alive, we can take a guess that someone named Brittany is somewhere between 19 and 25 years old in 2015 -- a pretty good guess considering we know nothing else about Brittany. The FiveThirtyEight team went on to analyze a bunch of other names and identify several interesting trends, but at the end of the article I wanted to look up my own name to see how well this method works for me. Since I was already analyzing the entire baby name database anyway, I decided to make a web tool out of FiveThirtyEight's methods and called it the Name Age Calculator: Enter your name and gender, and it will make a best guess at how old you are. Click here to look up your name on the Name Age Calculator I was a little disappointed to see that the method didn't work well for my name after all, but it nails the age of my father, who I was named after. So the Name Age Calculator indirectly got it right. The rise of the *aydens is pretty accurately captured: As is our recent fascination with Game of Thrones: There's plenty of other fun trends to explore. Have fun looking up your name, your family's names, and your friend's names. How close is it to guessing your name? If you'd like to tinker with the code or send in a bug report (or better yet, a bug fix!), I've open sourced the code on GitHub here. Caveats As with any data analysis project, there are some important caveats to this project. For one, previous research on the Social Security Administration baby name database suggests that records prior to 1940 are estimates and should be taken with a grain of salt. However, records after 1985 are based on actual Social Security records and are quite reliable. Another methodological caveat to this project is that the actuarial tables don't break down survivorship by name, and it's likely that people with certain names don't live as long as others due to various socioeconomic factors. There isn't much we can do about that, though, since the only input the web app receives is your name. --- ## The New York Times weather chart redux URL: https://www.randalolson.com/2015/08/13/the-new-york-times-weather-chart-redux/ Published: 2015-08-13 Categories: data visualization, python Tags: New York Times, reproducibility, weather Randy Olson shows how to recreate the famous New York Times weather charts in Python. One of my favorite pastimes is recreating and updating old New York Times graphics. It's great practice decomposing graphs into reproducible elements, and I always learn a ton about good graphic design in the process. If you're still learning data visualization yourself, I highly recommend doing the same. Last month, I updated the famous New York Times weather chart and was fortunate enough to contribute it to FiveThirtyEight. In this post, I'm going to briefly go over the process of how I made these charts. You can find all of the data and code for the FiveThirtyEight article on GitHub here. The first step is to find some up-to-date weather data. There's plenty of data sources out there for historical weather data, but I decided to focus on Wunderground.com because their historical weather pages are quite easy to parse. In the GitHub repository I linked above, there are two Python scripts: a scraper and a parser. The scraper downloads the web pages containing historical weather data for the weather station we point it at, while the parser uses BeautifulSoup to efficiently parse the HTML and transform the web pages into a flat CSV file. With the up-to-date weather data in hand, we can now turn to the design of the weather chart. The key point to reproducing the New York Times weather chart is realizing that it's just a bar chart, where every day is represented by three bars: A bar representing the record historical high and low temperatures for the day, a bar representing the average of the historical high and low temperatures for the day, and a bar representing the actual high and low temperatures for the day. Simple, right? All we have to do is draw three bars for every day, where the bar starts at the min and stops at the max. Since we have all of those values already, we can use a handful of lines of Python code to generate the chart. The Python code may look a bit complicated, but it's just a bunch of custom formatting for the chart. The core of the chart is generated with the three plt.bar() calls. To add something new to the weather chart, I added small red and blue dots to represent new (or tied) record high and low temperatures, respectively, for the year. These are small visual cues to quickly communicate when and how often the city is experiencing record temperatures without requiring much inspection from the user. Seattle, WA, for example, has had about 20 record hot days in the past year, which you can now quickly spot thanks to the numerous red dots. In the Python code, this is only a couple plt.scatter() calls for the days where the actual high/low matches the record high/low temperature. After that, I manually designed the legend in Omnigraffle since it was a bit tricky to generate programmatically. I'll leave that as a practice exercise if you're really looking to flex your matplotlib muscles. Voilà! We now have a reproducible method to generate arbitrary weather charts for any weather station on Wunderground. Enjoy! What are some other classic charts that you would like to see updated? --- ## Use the Baby Name Explorer to find out when your name was popular URL: https://www.randalolson.com/2015/08/13/use-the-baby-name-explorer-to-find-out-when-your-name-was-popular/ Published: 2015-08-13 Categories: data visualization Tags: baby names, Social Security Administration Randy Olson demonstrates a web app for exploring trends in American baby names. Around the same time I was working on the Name Age Calculator, I developed a simple tool to visualize trends in American baby names. Ever inventive, I named this web app the U.S. Baby Name Explorer. The idea behind the U.S. Baby Name Explorer is simple: Take the Social Security Administration baby name database and make visual interface for people to look up baby name trends. Whenever you enter a name and click Explore, it will show you the trends for both boys and girls back to the 1880s. There's some pretty humorous trends in the database, like the rise of the Tyrions thanks to Game of Thrones: and the rise of Nevaeh ("Heaven" backwards) after P.O.D. lead singer Sonny Sandoval named his daughter that in 2000. We sure follow some weird trends when naming our kids. Have fun exploring baby name trends. Here's a list of the 25 most gender neutral baby names if you're looking for some more interesting trends to look up. If you'd like to tinker with the code or send in a bug report (or better yet, a bug fix!), I've open sourced the code on GitHub here. Caveats Previous research on the Social Security Administration baby name database suggests that records prior to 1940 are estimates and should be taken with a grain of salt. However, records after 1985 are based on actual Social Security records and are quite reliable. --- ## Visualizing Indego bike share usage patterns in Philadelphia URL: https://www.randalolson.com/2015/07/18/visualizing-indego-bike-share-usage-patterns-in-philadelphia/ Published: 2015-07-18 Categories: analysis, data visualization Tags: bike share, Indego, machine learning, modeling, Philadelphia Randy Olson visualizes the Indego bike share usage patterns in Philadelphia to see how the locals are making use of the bike share program. One of the many things that I love about my new home town of Philadelphia is that the government openly shares curated data sets covering most of the governmental functions. Since I recently joined Philadelphia's Indego bike share program, I decided to start working with their bike usage data set to see what useful tools I could build. If you've ever used a bike share before, you know that one of the biggest fears is coming up to an empty bike share station when you need a bike. (Or similarly, coming up to a full station when you need to drop a bike off.) To help abate those fears, I've started monitoring the Indego bike usage API to see if I could model and predict when the bike share stations are most likely going to be empty or full. This tool is useful in two ways: It helps the Indego users by keeping them aware of the bad times to use certain bike share stations. It helps the Indego bike share service by predicting when a particular bike share station will require service (e.g., driving a truck out to pick up or drop off a large batch of bikes). Undoubtedly Indego's data science team is already performing some flavor of this model-and-predict scheme, but I thought it'd be fun to publicly tackle this problem and see how far I could get. For this post, I'll focus on visualizing patterns in the data, and will take a stab at prediction in a future post. Daily usage patterns One of the first steps toward building a model that can make any sort of useful prediction is to look at the existing patterns in the data. How are Philadelphians making use of the Indego bike share program? What does a typical day look like for the Indego bike share program? To get at those questions, I've been gathering the current status of each bike share station every 5 minutes since July 1, 2015. To provide some visuals of the data, I fit regressions to the usage patterns of each individual bike share station. The measure I'm using here to represent "station usage" is the percentage of a station's docks that are filled with bikes, where 100% represents a station full of bikes and 0% represents an empty station. keywords: Indego, bike share, Philadelphia, machine learning, modeling, visualization visualization description: Randy Olson visualizes the Indego bike share usage patterns in Philadelphia to see how the locals are making use of the bike share program. I've plotted each regression below, separated into three distinct categories: Outbound commuting stations: Stations where people take a bike from home to ride to work or school. Inbound commuting stations: Stations where people take a bike from work or school to ride back home. Underused stations: Stations that see minimal use compared to the other stations. As the two above plots show, many Philadelphians have adopted the Indego bike share program into their daily commute to work. Around 8 AM ET, we start to see bikes leaving several stations around town, which is followed shortly thereafter by an influx of bikes into stations at other parts of the city. Similarly around 5 PM ET, we see the reverse trend, with bikes heading back to the home stations. Unfortunately, several bike share stations seem to go mostly ignored. As shown in the above plot, these stations see little to no change in their usage throughout the day -- with the same bikes sitting in their docks day by day -- which perhaps means the stations need to be relocated. Mapping the daily usage patterns To provide a better spatial context to the above patterns, I mapped each bike share station onto an interactive map of Philadelphia and color-coded the stations by their usage pattern. As expected, the stations that see a large influx of bikes during work hours are in the primary business districts and education centers of Philadelphia. The bike stations along Market Street, around the University of Pennsylvania and Drexel University, and even up at Temple University are all places that Indego bike sharers ride to work. In contrast, most of the bike stations that see a decline in bikes in the morning are in residential areas further out in the city. This observation only piles on evidence that the Indego bike share program is being used for daily commutes to work and school moreso than joyrides by tourists. By this view, the Indego bike share program has been a resounding success so far. Some of the existing underused stations may require adjustment, but it's quite clear that Indego is here to stay. Weekly usage patterns Finally, I thought it would be interesting to show the weekly usage patterns of the stations. I've selected a handful of stations below and plotted their usage patterns, where darker red means "close to full of bikes" and darker blue means "close to empty." The stations at 11th & Reed and CHOP display the stereotypical commuting patterns that I discussed above. Interestingly, the CHOP station is one of a handful of stations that seems to be used almost exclusively for commuting, whereas most stations see some form of notable activity on the weekends. Above, I've visualized the weekly usage patterns of the station at 2nd & Germantown to highlight the irregular usage patterns of some of the Indego bike share stations. Even though the 2nd & Germantown station is used as an outbound commuter station on the weekdays, it's also quite popular as a station to reach the bars, restaurants, and activities in Northern Liberties on Friday night. At this point, I clearly need more data to properly model and predict the usage patterns since it's fairly clear that some bike stations are used differently at different times of the week. In the meantime... What else would help the Indego bike share program? I was previously thinking that we needed an Indego dock status tracker, but the most common devices are already covered: web | iOS | Android Do you have any ideas for what tools would be useful to supplement the Indego bike share program? Feel free to add your suggestions here in the comments. --- ## Rethinking the population pyramid URL: https://www.randalolson.com/2015/07/14/rethinking-the-population-pyramid/ Published: 2015-07-14 Categories: data visualization, tutorial Tags: critique, data visualization, population pyramid, U.S. Census Randy Olson critiques the popular population pyramid that the U.S. census likes to use so frequently to display population statistics. If you've ever browsed the U.S. Census population statistics pages, you've no doubt come across the famous population pyramid that they so frequently use to display the distribution of the U.S. population by age and gender. I was reading up about population pyramids last weekend and ran across an interesting quote that caught my eye: the use of a population pyramid is considered the best way to graphically illustrate the age and sex distribution of a given population. Now, I'm no expert at displaying population statistics, but I was shocked at this claim. Could it really be true that population pyramids are considered the best method for displaying population distributions? That line of thought ultimately led to the article below, where I raise three critiques of the population pyramid and present simpler and -- in my view -- more effective visualization methods. For this article, I used the 2010 U.S. Census population statistics, which you can find here in a machine-readable format. You can also find all of the code for these charts in my GitHub repository. Problems with the population pyramid 1) Violates the standard expectation of having the causal variable on the x-axis One of the most noticeable mistakes that the population pyramid makes is flipping the chart on its side to form a "pyramid" shape. I can only view this as an aesthetic flourish, since it violates one of the standard expectations of plotting: The causal variable should always be on the x-axis. When it comes to plotting, the x-axis is typically reserved for the independent variable, i.e., a fixed setting that has some sort of effect on another variable. In contrast, the y-axis is reserved for the dependent variable, i.e., the variable that shows some effect from varying the independent variable. The implication is that values on the x-axis cause some measurable effect on the values in the y-axis. This is why we always put the passage of time on the x-axis: it doesn't make sense to think of some other factor causing changes in the passage of time. (Until we discover time travel, anyway.) Since it doesn't make sense to think about a population's gender distributions having an effect on age -- and it makes far more sense to think about age having an effect on a population's gender distributions -- let's flip the axis of the pyramid so it's more in line with standard visualization practices. Now we don't have to reorient ourselves every time we look at the population pyramid, since the data is displayed more naturally. Ideally, the x-axis labels would be in between the "women" and "men" bars, but that was a bit tedious to pull off in my plotting software. Moving on... 2) Doesn't allow direct comparisons between the two categories The second flaw with population pyramids is that they make it difficult to compare the age distributions of men and women. For example, can you tell me at a glance if there's more men or women in the 25-29 age group? You'd have to look up the number of men and women in the 25-29 age group separately and make the comparison that way, when there's really no reason that the chart shouldn't be performing those comparisons for you. Let's rework the population pyramid to group the people by age, with separate bars for men and women. Now we have the exact same benefits of the population pyramid, with the additional benefit of being able to immediately discern whether there are more women or men in each age group. Arguably, we can now perform the same comparison between age groups as well -- for example, are there more 50-54-year-old women than 30-34-year-old men? -- but those comparisons become difficult the further the age groups are from each other. What's immediately apparent from this version of the population pyramid is: There are more young men than young women in the U.S., we reach gender parity around age 30, then men start dying out younger and leaving droves of widows behind starting at age 45. There's some really interesting implications in that data for the evolution of human sex ratios, but I'll leave that for another time. 3) Relative trends between the categories are masked by displaying absolute values There's clearly an interesting trend going on in the age 45+ groups where there are more women than men. But what's going on with the M:F ratio, especially in the 90+ categories? It's incredibly difficult to tell because these trends are masked when we display absolute values. If we're more interested in the relative trends between the two categories, we can drop the absolute values and instead show the percentage breakdown of the groups as I've done below. Now those trends I discussed above become abundantly clear, and we see that roughly 75% of U.S. adults aged 90+ are women. Sorry, straight men: your wife is probably going to outlive you. Of course, whether we would use this third chart solely depends on whether we care more about relative differences between the gender categories or the age distribution of the population. As with all charts, what data you should display depends on what story you want to tell with the data. In either case, I hope I've convinced you that the population pyramid -- as it's currently used -- is not quite ideal for telling either story. Lessons learned As with all of my long-winded articles critiquing a data visualization, I'll end with a brief summary of the main lessons we've learned. The causal variable (e.g., time or a parameter you control in an experiment) should always go on the x-axis. Group related data when within-group comparisons can be useful. What chart you use and what data you display depends on the story you want to tell. Don't try to force a story out of the wrong chart. Are there some other ways the population pyramid could be improved? Leave your suggestions in the comments. If you liked what you saw in this post and want to learn more, check out my Python data visualization video course that I made in collaboration with O'Reilly. In just one hour, I will cover these topics and much more, which will provide you with a strong starting point for your career in data visualization. --- ## 144 years of marriage and divorce in 1 chart URL: https://www.randalolson.com/2015/06/15/144-years-of-marriage-and-divorce-in-1-chart/ Published: 2015-06-15 Categories: data visualization Tags: divorce, marriage, usa Randy Olson charts out 144 years of marriage and divorce in the U.S. to see how the institution of marriage has evolved. I've always been curious about the history of marriage and divorce in the United States. We often hear about how divorce rates are in flux, or how marriage rates are declining, but we're rarely given a real sense of the long-term trends in marriage and divorce. Since I couldn't find a chart showing the long-term marriage and divorce trends in the U.S., I decided to crawl through the CDC's National Center for Health Statistics (NCHS) database and collect the data myself. If you're not familiar with the NCHS, they publish monthly reports on vital statistics about Americans such as birth rates, death rates, causes of death, and -- you guessed it -- marriage and divorce rates. Most of the historical data is hidden away in PDFs of these monthly reports, so I had the pleasure of scraping data from scans of dozens of CDC reports that were published 30 years before I was even born. I'll save you the laborious effort of repeating my scraping efforts by sharing the data set here. To provide a more visual view of the data set, I charted the per capita marriage and divorce rates below, with a few annotations to denote major historical events. Click here for the interactive version of the per capita chart It's fascinating to see the effects of WWI and WWII on marriage and divorce rates in the United States. At the beginning of America's entry into WWI (1917) and WWII (1941), we see notable spikes in marriage rates as the young conscripts rushed to the altar thinking it would be the last time they would see their lover. Similarly, after the conclusion of WWI (1918) and WWII (1945), those same young men and women coming back from the war seemed eager to elope and start a new life after spending years experiencing the destructive nature of war. Interestingly, the only notable spike in divorce rates in the past 144 years also followed the conclusion of WWII, likely due to many of the pre-WWII marriages coming to an abrupt end once the romance of wartime marriage wore off. The most notable drop in marriage rates occurred during The Great Depression in the early 1930s, with a sudden 25% drop in marriage rates during America's greatest time of hardship. It seems when Americans fall on hard times, marriage is one of the first things to take the back seat. One particularly confusing aspect of this data set was the fact that the post-war era in the 1950s and 1960s seemed to experience a significant drop in marriage rates, despite the fact that the 1950s and 1960s were known as a time of nearly-universal marriage in the U.S. To provide a clearer view of the 1950s and 1960s, I plotted the raw counts for marriages and divorces below. Click here for the interactive version of the raw counts chart With the raw counts in hand, the explanation for the drop in per capita marriage rates becomes abundantly clear: People weren't marrying less in the 1950s and 1960s, but the surge of newborn children during the Baby Boom artificially decreases the per capita rates. Once the Baby Boomers came of age in the 1970s, marriage rates returned to pre-WWII levels -- barring a slight drop in marriages during the dramatic conclusion of the Vietnam War (1975). Looking to more recent history, there has been a steady decline in marriage rates (and consequently, divorce rates) since the 1980s, with no sign of slowing down. In fact, when taking population into account, marriage rates in the U.S. are now at the lowest they've ever been in recorded U.S. history -- even lower than during The Great Depression! If you think you know why marriage rates have been declining in the U.S. since the 1980s, I'd be curious to hear your theories in the comments. --- ## Artificial Intelligence has crushed all human records in 2048. Here's how the AI pulled it off. URL: https://www.randalolson.com/2015/04/27/artificial-intelligence-has-crushed-all-human-records-in-2048-heres-how-the-ai-pulled-it-off/ Published: 2015-04-27 Categories: analysis, data visualization Tags: 2048, artificial intelligence, strategy, video games Randy Olson analyzes the AI that beat all of the human records in the video game "2048" to find out how it works. By now, we've all heard of the addictive tile-mashing game called 2048. Last week, I picked up 2048 for the first time and -- true to my nature -- I started designing an AI to beat the game for me the following day. It didn't take me long to find out that there's already some pretty good AIs out there, so I picked up the best 2048 AI I could find and fired several instances of it to see what it could do. Much to my surprise, it not only beat 2048... it crushed every human record in 2048 that I could find. Below is a video of the first 20 seconds of the AI hacking away at the game, mashing and merging tiles at superhuman speeds that we only wish we could match. Like a multi-armed genius, the AI played 1,000 games of 2048 simultaneously without a hitch. Some games ended in a few minutes due to a series of unfortunate random tile spawns, while others nearly lasted 4 hours and reached scores previously thought impossible. The worst instance achieved a score of 35,600, but even that instance managed to build the 2,048 tile and beat the game. Most instances ended with a score around 390,000 and a 16,384 tile, but the best instance built a 32,768 tile and stayed alive long enough to reach a score of 839,732. As far as I know, this is the highest score achieved in 2048 without undos. Perhaps even more impressive is how consistently the AI beats 2048. The AI reached the 2,048 tile -- and even the 4,096 tile -- in all 1,000 games, and reached the 16,384 tile in a large majority of them. In 1/3 of the games, the AI astonishingly reached the 32,768 tile, though it wasn't able to make it much further past that. (Though it's theoretically possible, if you're lucky.) For the rest of this post, I'll be looking at the game where the AI reached the high score of 839,732. In that instance, the AI beat the game in only 973 moves, which is about average for the AI. What's especially curious about the AI's progression is that it tends to reach tile X in about (X / 2) moves. For example, the AI reached the 16,384 tile in about 8,000 moves. At that rate, the AI would theoretically reach the 131,072 tile (the theoretically largest tile) in about 65,500 moves -- roughly 2x the number of moves it ended up lasting for -- if the random tile spawns played out in its favor. The AI's 2048 strategy To get a better understanding of how the AI managed to rack up such a high score, I analyzed its playing strategy on its highest-scoring game. Below, I'll outline 4 useful playing tips that the AI adopted to beat 2048. Some of these tips are fairly well-known -- and were even coded into the AI as heuristics -- but I figured it's good to cover the bases. Following these tips will undoubtedly help you improve your game and -- hopefully -- beat 2048. Ultimately, however, your survival toward the end of the game relies heavily on the random tile spawns working out in your favor; one poorly placed tile can spell the doom of your game. Tip #1: Keep your highest-value tile in one corner for the whole game One of the earliest strategies that players discovered for beating 2048 was to keep your highest-value tile in one of the corners for the entire game and slowly build it up. Some writers even called this a major design flaw of 2048 because it tends to make the game (relatively) easy to beat. It's no surprise, then, that the AI took advantage of this design flaw to beat 2048. In this game, the AI happened to choose the upper left corner, but all corners are equally viable. Pick one corner and stick to it. Tip #2: Keep your highest-value tiles lined up Another early strategy that 2048 players adopted was to maintain a row of monotonically increasing tiles as you build your main tile. I took a screenshot from my video above to demonstrate: Notice how the tiles are nicely lined up? The 64 is right next to the 32; the 32 is right next to the 16; and the 16 is right next to (what is about to become) the 8. When that line of 4s is combined into a 16, this configuration allows you to quickly compress the entire row into your next-highest tile and start again. Unsurprisingly, the AI adopted this strategy as well. The AI chose the top row as its primary row and kept 2nd-highest-value tile on the board was consistently right next its highest-value tile... ... and kept the 3rd-highest-value tile right next to the 2nd-highest-value tile in the top row. And so on. You'll also notice that the AI often had the 3rd-highest-value tile just under the highest-value tile as well, which I'm still trying to understand. Any thoughts? Tip #3: Keep the squares occupied Above, I mentioned that unfortunate random tile spawns can often spell the end of your game. One of the more interesting strategies that the AI seemed to adopt was to keep most of the squares occupied to reduce randomness and control where the tiles spawn. For most of the game, the AI maintained 12-15 (out of 16) tiles on the board and avoided merging too many tiles at once. While risky, this strategy ensures that tiles will spawn where you want them on the board. In this game, the AI kept the bottom-right corner of the board open so the low-value tiles would spawn around there, which it would then merge with nearby low-value tiles and move up the chain. If you choose a different corner, make sure to use the diagonally opposite corner as your "spawning ground." Tip #4: Maximize the number of possible merges on the board One of the mistakes that newer 2048 players tend to make is to try to merge everything quickly and leave as many squares open as possible. While such a strategy makes intuitive sense -- more open squares means you're less likely to get gridlocked -- focusing on immediately merging everything actually leads to shorter games. In its best game, 2/3 of the moves the AI made resulted in 2+ possible tiles that could be merged. At the extreme, some moves resulted in 6+ possible tile merges, but of course not all those merges could be made at once. Try playing a few games where you keep tiles lined up to merge -- but don't merge them until you have to -- and see if you last longer than usual. Have you beat this high score? If you've beaten this high score in 2048 (with an AI or otherwise) without undos, please let me know in the comments below! I'd love to hear how this AI can be beaten. --- ## Optimized walking tours of New York City and Philadelphia URL: https://www.randalolson.com/2015/03/25/optimized-walking-tours-of-new-york-city-and-philadelphia/ Published: 2015-03-25 Categories: analysis, data visualization, machine learning Tags: genetic algorithm, machine learning, optimization, traveling salesman problem, walking tour Randy Olson shows how machine learning can be used to optimize walking tours in New York City and Philadelphia. In the past couple weeks, I've shown you how to use machine learning to improve the efficiency of road trips all over the world, including the U.S., Europe, and South America. In this post, I'd like to show you how we can use the exact same algorithm to optimize walking tours in large cities. If you're anything like me whenever you stop in a big city, you have a few main attractions that you'd like to visit and end up haphazardly spending half your day staring at maps and trying to figure out the public transit system. However, by focusing on a limited few attractions, we're missing out on dozens of interesting places that we could've visited along the way. This is where the route optimization algorithm can come into play again: If we hand the algorithm the list of popular attractions in a city, it can provide us an efficient walking tour around the city that hits all of those attractions. Below, I've mapped out optimized walking tours for New York City and Philadelphia that hit the most popular attractions in each city according to TripAdvisor. These walking tours by no means represent the "definitive" walking tours for the cities, but both tours are guaranteed to give you a great taste of the cities while you're there. Optimized walking tour of New York City If skyscrapers, museums, and large crowds are your gig, then this walking tour is for you. In 4.5 hours of walking, you'll hit major sights like Central Park, Times Square, the Empire State Building, the Statue of Liberty, and more. This walking tour covers 14 miles of Manhattan, so come prepared with your walking shoes for what will add up to a half marathon. Of course, if you want to stop and actually enjoy the sights along the way (which I strongly advise!), this tour could easily end up taking several days. Between long lines, hour-long museum tours, and frequent stops for Cronuts, you'll likely only be able to hit a handful of stops per day before everything closes down. Keep that in mind when dividing this trip up over several days. Click here for an interactive version If 14 miles of walking sounds like too much, you can easily cut down on walking by taking the subway or bus for some of the longer legs of the trek. Or better yet, the same tour can be followed on a bike instead. Here's the Google Maps walking directions: [1] [2] [3] Here's the full list of attractions in order: American Museum of Natural History, Central Park West, New York, NY The Metropolitan Museum of Art, 1000 5th Avenue, New York, NY 10028 The Frick Collection, East 70th Street, New York, NY Central Park, New York, NY The Museum of Modern Art (MOMA), West 53rd Street, New York, NY St. Patrick's Cathedral, 5th Avenue, New York, NY Rockefeller Center, 45 Rockefeller Plaza, New York, NY 10111 Radio City Music Hall, Avenue of the Americas, New York, NY Theater District, New York, NY Times Square, New York, NY Bryant Park, New York, NY New York Public Library, New York, NY Chrysler Building, Lexington Avenue, New York, NY Grand Central Terminal, New York, NY The Morgan Library & Museum, Madison Avenue, New York, NY Empire State Building Observation Deck, 5th Avenue, New York, NY Madison Square Garden, Pennsylvania Plaza, New York, NY The High Line, New York, NY 10011 Chelsea Market, 9th Avenue, New York, NY Ground Zero Museum Workshop, West 14th Street, New York, NY Greenwich Village, New York, NY Washington Square Park, New York, NY SoHo, New York, NY Tenement Museum, Orchard Street, New York, NY St. Paul's Chapel, Broadway, New York, NY The National September 11 Memorial & Museum, New York, NY 10006 Statue of Liberty Optimized walking tour of Philadelphia The Liberty Bell, cheese steaks, and obnoxious sports fans: These are some of the things that Philadelphia is widely known for. But did you know that Philadelphia is home to one of the largest urban park systems in the U.S.? Or that it's home to America's first zoo? This 12-mile, 4-hour walking tour will take you through some of the best that Philadelphia has to offer, from beautiful parks to historic landmarks to breathtaking outdoor art collections. Again, don't take this tour lightly, and plan to spend at least a few days in total enjoying the sights. Click here for an interactive version The same disclaimers from the NYC walking tour apply here: If 12 miles of walking sounds like too much, you can easily cut down on walking by taking public transit for some of the longer legs of the trek. Or better yet, the same tour can be followed on a bike instead. Here's the Google Maps walking directions: [1] [2] [3] Here's the full list of attractions in order: Fairmount Park, Philadelphia, PA Please Touch Museum, 4231 Avenue of the Republic, Philadelphia, PA Philadelphia Zoo, West Girard Avenue, Philadelphia, PA Philadelphia Museum of Art, Benjamin Franklin Parkway, Philadelphia, PA Rocky Steps, Benjamin Franklin Parkway, Philadelphia, PA Eastern State Penitentiary, Fairmount Avenue, Philadelphia, PA Rodin Museum, Benjamin Franklin Parkway, Philadelphia, PA Barnes Foundation, 2025 Benjamin Franklin Parkway, Philadelphia, PA 19130 The Franklin Institute, North 20th Street, Philadelphia, PA Mutter Museum, South 22nd Street, Philadelphia, PA Penn Museum, 3260 South Street, Philadelphia, PA University of Pennsylvania, Philadelphia, PA Rittenhouse Square, Philadelphia, PA Cathedral Basilica of Saints Peter & Paul, Race Street, Philadelphia, PA Pennsylvania Academy Of The Fine Arts, North Broad Street, Philadelphia, PA Masonic Temple, North Broad Street, Philadelphia, PA Philadelphia City Hall, John F Kennedy Boulevard, Philadelphia, PA Reading Terminal Market, 51 N 12th St, Philadelphia, PA National Constitution Center, 525 Arch Street, Philadelphia, PA Fireman's Hall Museum, North 2nd Street, Philadelphia, PA Elfreth's Alley, Philadelphia, PA Christ Church, 20 North American Street, Philadelphia, PA National Museum of American Jewish History, Philadelphia, PA Independence National Historical Park, Philadelphia, PA Philadelphia's Magic Gardens, South Street, Philadelphia, PA Optimize your own walking tour Undoubtedly, you'd like to pick your own stops in your own cities and optimize a walking tour around that. Well, you're in luck! If you'd like to customize your own walking tour, I've released the Python code I used in this project with an open source license and instructions for how to optimize your custom walking tour. If Google Maps can route in the city you want to take walking tour in, then this code will work for you. You can find the code here. Happy exploring! --- ## Pure Michigan road trip, optimized URL: https://www.randalolson.com/2015/03/18/pure-michigan-road-trip-optimized/ Published: 2015-03-18 Categories: analysis, data visualization, machine learning Tags: genetic algorithm, machine learning, optimization, road trip, traveling salesman problem Randy Olson computes the optimal road trip to experience Pure Michigan. For the past couple weeks, I've been computing optimal road trips across the U.S. and Europe. This time, as a shout-out to my current home state, I made a road trip around Michigan. With the snow melting and a warm summer right around the corner, it's about time we start planning our summer fun. If there's one thing I know about Michiganders it's that they love their state more than anything, so I'm undoubtedly stepping on thin ice by picking only a handful of sites to visit in Michigan. Regardless, I thought I'd try my hand at creating the ultimate Pure Michigan experience by mapping out all of the Pure Michigan "hot spots." Below, I've mapped an optimized road trip circling the entire state of Michigan, hitting 43 sites along the way that are integral to the Pure Michigan experience. This trip makes a 2,098-mile (3,377 km) circle around the great state of Michigan, only spending 40 hours on the road so you can spend more time taking in the state's natural beauty. Feel free to start at any point on the circuit; as long as you follow the route in order from then on, you'll eventually hit all of the points and end up where you started. You may want to plan to take this trip sometime in the Fall so you can enjoy the Fall colors as you drive around the state. When you make your stop at St. Ignace, don't forget to hop on a ferry to enjoy a day of bike riding and eating fudge on Mackinac Island. I recommend checking out the Pure Michigan web site for advice on what to do at each stop. Here's the full list of stops in order: Ypsilanti, MI The Henry Ford Museum & Greenfield Village, Dearborn, MI Detroit, MI Port Sanilac, MI Flint, MI Frankenmuth, MI Bay City, MI Mt. Pleasant, MI East Tawas, MI Alpena, MI Gaylord, MI Cheboygan, MI Mackinaw City, MI St. Ignace, MI Mackinac Island, MI Sault Ste. Marie, MI Pictured Rocks National Lakeshore, Grand Marais, MI Marquette, MI Mt. Bohemia, Grant, MI Keweenaw Peninsula, Schoolcraft, MI Escanaba, MI Manistique, MI Harbor Springs, MI Petoskey, MI Charlevoix, MI Traverse City, MI Leelanau Peninsula, Leelanau, MI Sleeping Bear Dunes National Lakeshore, Empire, MI 49630 Cadillac, MI Manistee National Golf & Resort, Manistee, MI 49660 Ludington, MI Michigan's Adventure, Muskegon, MI Muskegon, MI Grand Rapids, MI Holland, MI South Haven, MI Kalamazoo, MI Binder Park Zoo, Battle Creek, MI Coldwater, MI Lansing, MI Jackson, MI Tecumseh, MI Ann Arbor, MI Make your own road trip If you'd like to customize your own road trip, I've released the Python code I used in this project with an open source license and instructions for how to optimize your custom road trip. You can find the code here. --- ## Computing the optimal road trip across South America URL: https://www.randalolson.com/2015/03/14/computing-the-optimal-road-trip-across-south-america/ Published: 2015-03-14 Categories: analysis, data visualization, machine learning Tags: genetic algorithm, machine learning, optimization, road trip, South America, traveling salesman problem Randy Olson shows you how to compute an epic road trip across South America. By popular request, I've created another follow-up to my posts about computing optimal road trips across the U.S. and Europe. This time, I made an optimal road trip around South America. If you'd like to get into the nitty-gritty of how these road trips are created, check out the first post about the U.S. road trip. South America is yet another massive and diverse continent, so again there's no way I'll be able to pick a series of stops that will please everyone. However, by combining recommendations from TripAdvisor and Huffington Post, I was able to create a trip with a nice mix between beautiful outdoor sights and lively cities across the entire South American continent. Below is the optimized route between those stops. This trip ended up mostly following along the coastline of South America, covering roughly 18,148 miles (29,206 km) of driving, ferrying, and flying. You should plan to spend 3 or more months on this trip if you want to really enjoy it. Google Maps isn't routing between Colombia and its neighboring countries, so I drew direct lines between Quito Bogota & Santa Marta Isla Margarita that may need to be taken by plane instead. Your best bet with this trip is to start at Isla Margarita, Venezuela and follow the route south from there. Here's the full list of stops in order: Isla Margarita, Venezuela Angel Falls, Venezuela Jericoacoara, Brazil Fortaleza, Brazil Natal, Brazil Salvador, Brazil Trancoso, Brazil Belo Horizonte, Brazil Buzios, Brazil Rio de Janeiro, Brazil Sao Paulo, Brazil Curitiba, Brazil Florianopolis, Brazil Gramado, Brazil Punta del Este, Uruguay Asuncion, Paraguay Buenos Aires, Argentina Torres del Paine, Chile El Calafate, Argentina San Carlos de Bariloche, Argentina Santa Cruz, Chile Santiago, Chile Vina del Mar, Chile San Pedro de Atacama, Chile Nuestra Senora de La Paz, Bolivia Lake Titicaca, Bolivia/Peru Arequipa, Peru Urubamba, Peru Cusco, Peru Lima, Peru Mancora, Peru Quito, Ecuador Bogota, Colombia Cartagena, Colombia Santa Marta, Colombia Taganga, Colombia Make your own road trip If you'd like to customize your own road trip, I've released the Python code I used in this project with an open source license and instructions for how to optimize your custom road trip. You can find the code here. Happy exploring! --- ## Computing the optimal road trip across Europe URL: https://www.randalolson.com/2015/03/10/computing-the-optimal-road-trip-across-europe/ Published: 2015-03-10 Categories: analysis, data visualization, machine learning Tags: Europe, genetic algorithm, machine learning, optimization, road trip, traveling salesman problem Randy Olson shows you how to compute an epic road trip across Europe. As a follow-up to my post about computing optimal road trips across the U.S., I thought it'd be fun to make a road trip map for Europe as well. If you'd like to get into the nitty-gritty of how these road trips are created, check out the first post. Europe is a massive continent with a diverse mix of cultures, so there's no way I'm going to pick a set of stops that will please everyone. However, last year Business Insider published a nice article listing "50 Places In Europe You Need To Visit In Your Lifetime." These stops seemed like a nice mix of inner city exploration and outdoorsy fun from an eclectic collection of countries, so I made a map showing what an epic road trip across Europe hitting most of these spots would look like. (Some of the locations couldn't be reached by car, so I had to exclude them.) In total, the trip covers 16,287 miles (26,211 km) and 14 days of driving, so expect to dedicate at least 3 months if you're going to take on this road trip. You may notice that eastern Europe is somewhat underrepresented in this road trip, so if you want the full taste of Europe, it'll be worthwhile to add some stops between Poland and Estonia. Here's the full list of stops in order: Innsbruck, Austria Munich, Germany Pag, Croatia Venice, Italy Tuscany, Italy Florence, Italy Rome, Italy Vatican City Amalfi, Italy Gozo, Malta Dubrovnik, Croatia Santorini, Thira, Greece Rila Monastery, Rilski manastir, Bulgaria Istanbul, Turkey Sighisoara, Mures County, Romania Budapest, Hungary Vienna, Austria Prague, Czech Republic Krakow, Poland Jägala Waterfall, 74205 Harjumaa, Estonia Lapland, Finland ICEBAR, Marknadsvägen, Jukkasjärvi, Sweden Bergen, Norway Copenhagen, Denmark Berlin, Germany Amsterdam, Netherlands Keukenhof, Stationsweg, Lisse, Netherlands Edinburgh, United Kingdom Inverness, United Kingdom Ballybunion, Ireland Cliffs of Moher, Clare, Ireland Cornwall, England Stonehenge, Amesbury, United Kingdom London, United Kingdom Brussels, Belgium Paris, France Pamplona, Spain Lagos, Portugal Granada, Spain Ibiza, Spain Barcelona, Spain Luberone, Bonnieux, France Nice, France Monte Carlo, Monaco Interlaken, Switzerland Make your own road trip If you'd like to customize your own road trip, I've released the Python code I used in this project with an open source license and instructions for how to optimize your custom road trip. You can find the code here. Happy road tripping! --- ## Computing the optimal road trip across the U.S. URL: https://www.randalolson.com/2015/03/08/computing-the-optimal-road-trip-across-the-u-s/ Published: 2015-03-08 Categories: analysis, data visualization, machine learning Tags: genetic algorithm, machine learning, optimization, road trip, traveling salesman problem Randy Olson shows you how to compute an epic road trip across the U.S. Last week, Tracy Staedter proposed an interesting idea to me: Why not use the same algorithm from my Where's Waldo article to compute the optimal road trip across every state in the U.S.? Visiting every U.S. state has long been on my bucket list, so I jumped on the opportunity and opened up my machine learning tool box for another quick weekend project. Note: If you're not interested in the technical details of the project, skip down to the Road trip stopping at major U.S. landmarks section. Planning the road trip One of the hardest parts of planning a road trip is deciding where to stop along the way. Given how large and diverse the U.S. is, it's especially difficult to make a road trip that will appeal to everyone. To stand a chance at making an interesting road trip, Tracy and I laid out a few rules from the beginning: The trip must make at least one stop in all 48 states in the contiguous U.S. The trip would only make stops at National Natural Landmarks, National Historic Sites, National Parks, or National Monuments. The trip must be taken by car and never leave the U.S. With those objectives in mind, Tracy compiled a list of 50 major U.S. landmarks -- one in each state excluding Alaska/Hawaii and including D.C., and two in California. Tracy wrote about that process on Discovery News here. The result was an epic itinerary with a mix of inner city exploration, must-see historical sites, and beautiful natural landscapes. All that was left was to figure out the path that would minimize our time spent driving and maximize our time spent enjoying the landmarks. Image credit: Dean Franklin Computing the optimal road trip across the U.S. With the list of landmarks in hand, the next step was to find the "true" distance between all of the landmarks by car. Since we can't just drive a straight line between every landmark -- driving by car has this pesky limitation of having to stay on roads -- we needed to find the shortest route by road between every landmark. If you've ever used Google Maps to get the directions between two addresses, that's basically what we had to do here. Except this time, we needed to look up 2,450 directions to get the "true" distance between all 50 landmarks -- a monumental task if we had to do it by hand. Thankfully, the Google Maps API makes this information freely available, so all it took was a short Python script to calculate the distance and time driven for all 2,450 routes between the 50 landmarks. Now with the 2,450 landmark-landmark distances, our next step was to approach the task as a traveling salesman problem: We needed to order the list of landmarks such that the total distance traveled between them is as small as possible if we visited them in order. This means finding the route that backtracks as little as possible, which is especially difficult when visiting Florida and the Northeast. If you read my Where's Waldo article, you're already aware of how difficult it can be to solve route optimization problems like this one. With 50 landmarks to put in order, we would have to exhaustively evaluate 3 x 1064 possible routes to find the shortest one. To provide some context: If you started computing this problem on your home computer right now, you'd find the optimal route in about 9.64 x 1052 years -- long after the Sun has entered its red giant phase and devoured the Earth. This complication is why Google Map's route optimization service only optimizes routes of up 10 waypoints, and the best free route optimization service only optimizes 20 waypoints unless you pay them a lot of money to dedicate some bigger computers to it. The traveling salesman problem is so notoriously difficult to solve that even xkcd poked fun at it: Clearly, we need a smarter solution if we want to take this epic road trip in our lifetime. Thankfully, the traveling salesman problem has been well-studied over the years and there are many ways for us to solve it in a reasonable amount of time. If we're willing to accept that we don't need the absolute best route between all of the landmarks, then we can turn to smarter techniques such as genetic algorithms to find a solution that's good enough for our purposes. Instead of exhaustively looking at every possible solution, genetic algorithms start with a handful of random solutions and continually tinkers with these solutions -- always trying something slightly different from the current solutions and keeping the best ones -- until they can't find a better solution any more. I've included a visualization of a genetic algorithm solving a similar routing problem below. Road trip stopping at major U.S. landmarks After less than a minute, the genetic algorithm reached a near-perfect solution that makes a complete trip around the U.S. in only 13,699 miles (22,046 km) of driving. I've mapped that route below. Note: There's an extra stop in Cleveland to force the route between Vermont and Michigan to stay in the U.S. rather than go through Canada. If you're able to drive through Canada without issue, then take the direct route through Canada instead. Here's the Google Maps of the route: [1] [2] [3] [4] [5] [6] (Note that Google maps itself only allows 10 waypoints to be routed at a time, hence why there's multiple Maps links.) Assuming no traffic, this road trip will take about 224 hours (9.33 days) of driving in total, so it's truly an epic undertaking that will take at least 2-3 months to complete. The best part is that this road trip is designed so that you can start anywhere on the route as long as you follow it from then on. You'll hit every major area in the U.S. on this trip, and as an added bonus, you won't spend too long driving through the endless corn fields of Nebraska. Here's the full list of landmarks in order: Grand Canyon, AZ Bryce Canyon National Park, UT Craters of the Moon National Monument, ID Yellowstone National Park, WY Pikes Peak, CO Carlsbad Caverns National Park, NM The Alamo, TX The Platt Historic District, OK Toltec Mounds, AR Elvis Presley's Graceland, TN Vicksburg National Military Park, MS French Quarter, New Orleans, LA USS Alabama, AL Cape Canaveral Air Force Station, FL Okefenokee Swamp Park, GA Fort Sumter National Monument, SC Lost World Caverns, WV Wright Brothers National Memorial Visitor Center, NC Mount Vernon, VA White House, Washington, DC Colonial Annapolis Historic District, MD New Castle Historic District, Delaware Cape May Historic District, NJ Liberty Bell, PA Statue of Liberty, NY The Mark Twain House & Museum, CT The Breakers, RI USS Constitution, MA Acadia National Park, ME Mount Washington Hotel, NH Shelburne Farms, VT Fox Theater, Detroit, MI Spring Grove Cemetery, OH Mammoth Cave National Park, KY West Baden Springs Hotel, IN Abraham Lincoln's Home, IL Gateway Arch, MO C. W. Parker Carousel Museum, KS Terrace Hill Governor's Mansion, IA Taliesin, WI Fort Snelling, MN Ashfall Fossil Bed, NE Mount Rushmore, SD Fort Union Trading Post, ND Glacier National Park, MT Hanford Site, WA Columbia River Highway, OR San Francisco Cable Cars, CA San Andreas Fault, CA Hoover Dam, NV Bonus: Road trip stopping at popular U.S. cities If you're more of a city slicker, the road trip above may not look very appealing to you because it involves spending a lot of time outdoors. But worry not, for I created a second road trip just for you! The road trip below stops at the TripAdvisor-rated Best City to Visit in every contiguous U.S. state. Note: Again, there's an extra stop in Cleveland to force the route between New Hampshire and Michigan to stay in the U.S. rather than go through Canada. If you're able to drive through Canada without issue, then take the direct route through Canada instead. But really, Cleveland is a nice city to stop in (ranked #53 on TripAdvisor). Here's the Google Maps of the route: [1] [2] [3] [4] [5] [6] This road trip will more-or-less follow the same path as the major U.S. landmarks trip, covering a slightly shorter 12,290 mile (19,780 km) route around the U.S. Some larger states -- like California and Texas -- may have multiple cities you'd like to visit, so it's probably worthwhile to stop at other larger cities along the route. You may note that cities from North Dakota, Vermont, and West Virginia are missing. Out of the top 400 recommended cities to visit on TripAdvisor, none were from North Dakota, Vermont, nor West Virginia. This is especially interesting because TripAdvisor reviewers recommend cities like Flint, MI -- the 7th most crime-ridden city in the U.S. -- over any city in North Dakota, Vermont, and West Virginia. I'll leave the interpretation of that fact to the reader. Here's the full list of cities in order: Oklahoma City, Oklahoma Wichita, Kansas Denver, Colorado Albuquerque, New Mexico Phoenix, Arizona Las Vegas, Nevada San Francisco, California Portland, Oregon Seattle, Washington Boise, Idaho Park City, Utah Jackson, Wyoming Billings, Montana Sioux Falls, South Dakota Omaha, Nebraska Des Moines, Iowa Minneapolis, Minnesota Milwaukee, Wisconsin Chicago, Illinois Indianapolis, Indiana Louisville, Kentucky Columbus, Ohio Detroit, Michigan Cleveland, Ohio Manchester, New Hampshire Portland, Maine Boston, Massachusetts Providence, Rhode Island New Haven, Connecticut New York City, New York Ocean City, New Jersey Philadelphia, Pennsylvania Wilmington, Delaware Baltimore, Maryland Washington, D.C. Virginia Beach, Virginia Charlotte, North Carolina Charleston, South Carolina Orlando, Florida Atlanta, Georgia Nashville, Tennessee Birmingham, Alabama Jackson, Mississippi New Orleans, Louisiana Houston, Texas Little Rock, Arkansas Branson, Missouri Make your own road trip If you'd like to customize your own road trip, I've released the Python code I used in this project with an open source license and instructions for how to optimize your custom road trip. You can find the code here. What about other parts of the world? I've made another version for Europe here and for South America here. I also made a road trip for Michigan, and optimized walking tours for NYC and Philadelphia. Check 'em out! Conclusions The saying goes, "A journey of a thousand miles begins with a single step." Really, that's not true. Every major journey begins with a plan: where you're going, where you're stopping along the way, and how you're getting there. I hope this article convinced you that machine learning can play a crucial role in that planning phase and save you a ton of time along the way. Of course, it may not be practical for you to take a road trip of epic proportions like the ones described here. But really, this algorithm works just as well when you're planning a smaller trip within your state as when you're planning a larger trip spanning the entire world. All the algorithm needs are the distances travelled between every stop so it can try to compute the optimal route. How you get between those stops is up to you. Happy road tripping! --- ## Revisiting the Six Degrees of Kevin Bacon URL: https://www.randalolson.com/2015/03/04/revisiting-the-six-degrees-of-kevin-bacon/ Published: 2015-03-04 Categories: analysis, data visualization Tags: Bacon number, Eric Roberts, film, Kevin Bacon, six degrees of separation Randy Olson revisits the Six Degrees of Kevin Bacon to see if Kevin is really the Center of the Hollywood Universe. In early 1994, three Albright College students were watching Footloose during a heavy snowstorm. By pure coincidence, the next movie that came on the television was The Air Up There, another popular film starring Kevin Bacon. Intrigued by the coincidence, the students started counting how many films Kevin Bacon had acted in and speculating how many actors he had appeared on film with. It didn't take long for the trio to turn their interest into a game, trying to link every actor to Kevin Bacon through a series of shared movie appearances. What began as an inside joke quickly spread across the nation and became a popular parlor game. Players may ponder: What series of films do you think Emma Watson and Kevin Bacon are linked through? You may be surprised: John Cleese links Watson with Bacon directly, leaving Watson with a "Bacon number" of 2 (the number of shared movie appearances linking Watson and Bacon). This game is now widely known as the Six Degrees of Kevin Bacon -- named so because the Albright College trio had found that no actor is more than 6 shared movie appearances ("degrees") away from indirectly collaborating with Kevin Bacon. Interestingly enough, the trio had stumbled across one of the first examples of the small-world phenomenon: Because everyone is so widely connected in modern times, we're connected to every other human in the world by no more than 6 links of mutual friends. Apparently, the same rules apply to actor collaboration networks. It's been over two decades since the Six Degrees of Kevin Bacon was invented. We now have a massive database of movies and vastly more powerful computers to look at this problem. It's about time we revisit the Six Degrees of Kevin Bacon to see if the trio's findings hold up. The Six Degrees of Kevin Bacon The Albright College trio picked Kevin Bacon because he's been a highly productive actor, appearing in over 60 films with a wide range of actors. According to the aptly-named Oracle of Bacon, Kevin Bacon has appeared in film with 3,031 actors during his film career. When the Oracle of Bacon calculated the Bacon number for 1.91 million other actors, the trio's finding were validated: 99%+ of all actors have a Bacon number of 5 or less. Below, I charted distribution of Bacon numbers for all 1.91 million actors. Incredibly, there is a single actor out there who starred in such an obscure film that he/she has a Bacon number of 10. Props to you if you find out who it is. Alongside the Bacon numbers, I charted the distribution of "Roberts" numbers. Eric Roberts stands out as an actor who has had such a productive movie career that he has 379 movie and TV credits to his name as of March 2015. It's no surprise, then, that Eric Roberts has appeared on the silver screen with 8,398 other actors -- truly an unprecedented accomplishment in Hollywood. Even more interesting is that Kevin Bacon actually isn't the "Center of the Hollywood Universe" as so many have claimed. When the Oracle of Bacon calculated the average number of movie appearances connecting every actor to every other actor, they found that Kevin Bacon didn't even make the top 100 most "central" actor list. In fact, the top 10 most "central" actors are: Eric Roberts Michael Madsen Danny Trejo Samuel L. Jackson Harvey Keitel Robert De Niro Willem Dafoe Malcolm McDowell Donald Sutherland Michael Caine That's right. You'd never guess it, but Eric Roberts -- the movie villain who we're always happy to see killed at the end of the movie -- is the real Center of the Hollywood Universe. The Three Degrees of Eric Roberts The most impressive part about Roberts' career is how few links it takes to connect him to any other actor out there. Shown in the chart below, 1 in 4 of the 1.91 million actors on IMDb are within 2 degrees of separation from Eric Roberts. In contrast, fewer than 1 in 5 are within 2 degrees of separation from Kevin Bacon. By 3 degrees out, Eric Roberts can be connected to 88% of all actors out there, which includes nearly all of the well-known actors. At this point, it's amazing that Kevin Bacon and Eric Roberts never starred in a movie together! My point? We should rename the Six Degrees of Kevin Bacon to the Three Degrees of Eric Roberts -- or at least the Three Degrees of Kevin Bacon. Given that Kevin Bacon dislikes the Six Degrees game, I'm sure he would see this as a welcome change. If you'd like to read more about how the Oracle of Bacon works, check out their How It Works page. If you find yourself stumped when playing the Three Degrees of Eric Roberts, the Oracle of Bacon also provides a lookup service. Now get out there and impress your friends with your impressive movie knowledge! --- ## Top 25 richest living comedians URL: https://www.randalolson.com/2015/03/04/top-25-richest-living-comedians/ Published: 2015-03-04 Categories: data visualization Tags: actors, comedians, rankings, wealth Randy Olson charts the top 25 richest living comedians to see how they made their riches doing comedy. It's fairly common knowledge that comedy isn't a terribly lucrative career. Not only do most comedians spend decades doing small-time standup hoping to be discovered, but most of those comedians never end up being discovered either. But what about the comedians that did hit it big? To provide some insight into what it takes to be a successful -- and rich -- comedian, I charted the top 25 richest living comedians below. The first thing you'll notice is that only one comedian made it into the top 25 by doing standup. Terry Fator, the lovable ventriloquist who won America's Got Talent in 2007. After winning America's Got Talent, Fator went on to land a series of lucrative gigs in Las Vegas and earned himself standup contracts worth millions. The rest of the comedians made their millions either by writing for TV shows or starring in TV shows and film. Perhaps the biggest success story on this list is the TV show Seinfeld, which made Jerry Seinfeld and co-writer Larry David rich beyond imagination when it was sold into syndication. Matt Groening (The Simpsons & Futurama), Trey Parker/Matt Stone (South Park), and Seth MacFarlane (Family Guy & American Dad) also pack this list with wildly successful comedy cartoon shows, showing us just how profitable running a TV show is nowadays. Then there's the late night comedy show hosts -- David Letterman and Jay Leno -- who rake in absurd salaries for making us laugh every night. (David Letterman brings in $20 million per year!) Not to be forgotten are the comedy film producers and actors such as Adam Sandler, Jim Carrey, and Bill Murray. Adam Sandler has been the most successful comedian in the film business, banking over $62 million from the film Anger Management alone. It's no wonder Sandler keeps pumping out new films every year even while his ratings are dropping. So, what can we take from this list? It looks like landing a comedy TV show on a major network is the true gold mine for comedians. Standup acts, for the most part, aren't where the real money is, but they're a great way to get your name out there when you're just getting started. --- ## Design critique: Putting Big Pharma spending in perspective URL: https://www.randalolson.com/2015/03/01/design-critique-putting-big-pharma-spending-in-perspective/ Published: 2015-03-01 Categories: data visualization Tags: big pharma, budget, design critique, marketing, r&d, remake Randy Olson critiques a recently popular infographic on the odd spending habits of Big Pharma companies. Recently, /r/DataIsBeautiful began hosting weekly visualization redesign competitions challenging everyone to come up with better and less misleading designs of existing graphics. Below is the first of many redesigns and detailed critiques that I will be working on. In early 2015, John Oliver and his team released an excellent exposé on Big Pharma and their shady marketing tactics. Shortly thereafter, Leon Markovitz from dadaviz released the following bubble chart to feed the ensuing anti-Big Pharma news cycle. While this bubble chart visualizes an interesting phenomenon, several aspects of the chart can be improved to tell a more accurate and complete story. Below, I will outline and address three such improvements. Better design through better chart choices It has been shown time and time again that circles are terrible for making comparisons. Worse, the above bubble chart makes it even more difficult to use the circles to compare values by not even overlapping them. The visual representations of the data in this graphic are minimally useful, and most readers will simply rely on reading and comparing the numbers inside each circle, which renders the chart a fancy-looking data table. Most basic guides for selecting charts recommend the use of bar charts for comparisons of data. Let's rework the above chart into a bar chart. The bar chart works much better than the bubble chart for comparing the company's marketing and R&D budgets by placing them on the same axis. It also still allows the viewer to look up approximate budget numbers via the x-axis grid lines. However, the bar chart is also quite cluttered because it's comparing 2 values for 10 companies. This is where it's important to think about the purpose of the chart. Leon wanted to use this chart to communicate the fact that "9 out of 10 Big Pharma companies spend more on marketing than R&D." This fact can more effectively be communicated by a scatter plot, as I've demonstrated below. In the chart below, each square represents a company. The key to this scatter plot is the line running diagonally through the center of the chart, which represents parity between marketing and R&D spending. Now the viewer can immediately tell how many Big Pharma companies spend more on marketing than R&D: They need only count the number of squares above and below the line of parity. As an added advantage, the scatter plot still allows the viewer to gauge approximate budgets for each company and allows for a third dimension of data -- total company revenue in this case -- to be visualized via the size of the squares. Even though the identity of the individual companies are lost in the scatter plot, this issue could be remedied by annotating the graph with the names, changing the squares to pictures of the company logos, or even turning the graphic into an interactive. I did not do so here because the company names are not particularly important to the story. Don't forget to normalize Another basic mistake in the original bubble chart was that the data was not normalized in any way, making comparisons between the Big Pharma companies precarious. Taken at face value, the non-normalized numbers seem to indicate that Johnson & Johnson is a marketing giant and far more invested in marketing than Astra Zeneca. These numbers completely ignore the fact that Johnson & Johnson brings in far more revenues than Astra Zeneca; when we take both company's total revenues into account, Astra Zeneca actually spends a higher percentage of its revenues (28%) on marketing than Johnson & Johnson (24%). Below, I normalized all of the expenditures by each company's 2013 yearly revenues. By normalizing the expenditures, the graph now tells a more complete story: We can meaningfully compare the Big Pharma companies and see that most of them spend about 15% of their revenues on R&D and 20-25% of their revenues on marketing, with Roche and Eli Lilly & Co. being the odd ones out sitting on the line of parity. Provide meaningful context Perhaps the most egregious oversight in the design of the original bubble chart was the failure to provide any meaningful context to the data. The viewer was left with the fact that "9 out of 10 Big Pharma companies spend more on marketing than R&D," but many viewers don't know if a large marketing budget is normal for a company or not. Left to their own devices, many viewers (especially those who watched John Oliver exposé) assumed "R&D good, marketing BAD" and immediately grabbed their pitchforks and aimed them at Big Pharma. To provide at least some context to the data, I looked up the 2013 marketing and R&D budgets of 6 large companies and plotted them alongside the Big Pharma companies. The companies are: Samsung Intel Microsoft Google Toyota General Motors These companies were picked based on the ease of looking up their budget and revenue information. Unsurprisingly, not all companies make this information readily accessible on the internet. At least based on the companies chosen, it appears that Big Pharma as a whole is an outlier when it comes to marketing budgets. Even Samsung with its infamous $14bn marketing budget only spends ~8% of its revenues on marketing. The only company that even comes close to Big Pharma in terms of marketing is Intel, but it still spends more on R&D than marketing. Perhaps the pitchforks over Big Pharma's apparently overgrown marketing budget were warranted, but we didn't know until at least some context was provided. Conclusions Well-designed data visualizations are one of the most effective mediums for communicating information today. We must be careful when designing visualizations to make sure that they tell the whole truth rather than bend statistics to tell the story we want to hear. In this critique, I have covered 3 common oversights that lead to bad and/or misleading visualizations: Selection of a proper chart Normalizing data Providing meaningful context Before sharing your visualizations in the future, please be sure to review your work to ensure that you didn't hit one of these common pitfalls. If you liked what you saw in this post and want to learn more, check out my Python data visualization video course that I made in collaboration with O'Reilly. In just one hour, I will cover these topics and much more, which will provide you with a strong starting point for your career in data visualization. --- ## Here's Waldo: Computing the optimal search strategy for finding Waldo URL: https://www.randalolson.com/2015/02/03/heres-waldo-computing-the-optimal-search-strategy-for-finding-waldo/ Published: 2015-02-03 Categories: analysis, data visualization, machine learning Tags: genetic algorithm, machine learning, optimization, traveling salesman problem, where's waldo Randy Olson uses machine learning techniques to find the best search strategy for finding Waldo. As I found myself unexpectedly snowed in this weekend, I decided to take on a weekend project for fun. While searching for something to catch my fancy, I ran across an old Slate article claiming that they found a foolproof strategy for finding Waldo in the classic "Where's Waldo?" book series. Now, I'm no Waldo-spotting expert, but even I could tell that the strategy they proposed there is far from perfect. That's when I decided what my weekend project would be: I was going to pull out every machine learning trick in my tool box to compute the optimal search strategy for finding Waldo. I was going to crush Slate's supposed foolproof strategy and carve a trail of defeated Waldo-searchers in my wake. "But Randy, don't you have better things to work on? You know, curing cancer, solving world hunger... ANYTHING else?", a sane person would have said at that point. Too bad that sane person wasn't around. What is "Where's Waldo"? For the poor souls who have no clue who Waldo is, I'll defer to the Wikipedia description: "Where's Waldo?" is a series of children's books created by English illustrator Martin Handford. The books consist of a series of detailed double-page spread illustrations depicting dozens or more people doing a variety of amusing things at a given location. Readers are challenged to find a character named [Waldo] hidden in the group. [Waldo's] distinctive red-and-white-striped shirt, bobble hat, and glasses make him slightly easier to recognize, but many illustrations contain "red herrings" involving deceptive use of red-and-white striped objects. Here's an example of a classic "Where's Waldo?" illustration: Here's Waldo Thankfully, the Slate article provided a chart that made it dead easy to acquire all 68 of Waldo's coordinates in the primary 7 editions of the "Where's Waldo?" books. I've reproduced those coordinates visually below. You can download the data file here. If we perform a kernel density estimation of these points, we start to see some interesting trends already: Waldo almost never appears in the top left corner. That's because there was always some postcard from Waldo in the top left corner describing the setting and some interesting facts about it. Waldo is rarely located on the edges. Slate's Ben Blatt hypothesized that this was done on purpose because the edges are "locations that might be construed as too obvious" and are "where children and adults alike might begin their search." Waldo is never located on the very bottom of the right page. I was unsure about the reason for this at first, but Chris Metzger offered a probable explanation: Whenever you flip to the next page in a book, the bottom of the right page is the first thing you see. Thus, the bottom of the right page would be one of the worst places to hide Waldo because that's the most-viewed part of the book. Computing the optimal search strategy Now on to the real fun! I decided to approach this problem as a traveling salesman problem: We need to check every possible location that Waldo could be at while taking as little time as possible. That means we need to cover as much ground as possible without any backtracking. In computer terms, that means we're making a list of all 68 points that Waldo could be at, then sorting them based on the order that we're going to visit them. So now we just need to try every possible arrangement of the points and find the one with the shortest distance traveled. Easy, right? Wrong. Those 68 points can be arranged in ~2.48 x 1096 possible ways. To provide some context, that's more possible arrangements than the number of atoms in the universe. That's so many possible arrangements that even if finding Waldo became an international priority and the world banded together to dedicate the 8.25 million computing cores from the world's 10 largest supercomputers to the job, it would still take ~9.53 x 1077 years -- about 6.35 x 1067x longer than the universe has existed -- to exhaustively evaluate all possible combinations. (Generously assuming that each core could perform 10,000 evaluations per second.) In other words: if we don't have a smarter solution, Waldo is as gone as Carmen Sandiego. Thankfully, there are plenty of smarter methods for approximating the optimal search path for finding Waldo. Below, I visualized the best solution over time of one such method -- a genetic algorithm -- that found a nearly-perfect solution. As you can see, genetic algorithms continually tinker with the solution -- always trying something slightly different from the current best solution and keeping the better one -- until they can't find a better solution any more. (Note: Because genetic algorithms -- like many optimization algorithms -- are stochastic in nature, they won't always result in the exact same solution at the end.) After running the genetic algorithm for about 5 minutes, I ended up with the solution below. I colored the paths by whether they're in the first (blue), second (orange), third (green), or final (red) 1/4 of the path. This path represents one of the shortest possible paths to follow on the page to find Waldo, so if we followed this path exactly, we'd most likely find Waldo much faster than someone following a more basic technique. If you'd like to learn more about the methodology used here, see the accompanying methods and code document. (For those interested: I also tried a standard hillclimber algorithm, but it always converged on a worse solution than the genetic algorithm.) Of course, we should never take results from machine learning too literally. A robot might be able to follow this path perfectly, but I wouldn't be able to remember that path unless it was etched on every page for me. Instead, I think we can take some general lessons from the path that the genetic algorithm discovered: The bottom of the left page is a good place to start. If Waldo isn't on the bottom half of the left page, then he's probably not on the left page at all. The upper quarter of the right page is the next best place to look. Waldo seems to prefer to hide on the upper quarter of the right page. Next check the bottom right half of the right page. Waldo also has an aversion to the bottom left half of the right page. Don't bother looking there until you've exhausted the other hot spots. I annotated the best solution with a general path to follow when searching for Waldo. If you don't find Waldo at the end of that trail, then you've got an outlier and should check the middle of the pages or the top left and right. How does this strategy compare? Unfortunately, I lost my old copies of "Where's Waldo?" ages ago in a move, so I couldn't test it out for myself. I'd love to put this strategy to the test, though, to see how much faster it is than the Slate strategy. The U.S. publishers of "Where's Waldo?" generously sent me the entire series of books, so I was finally able to put this strategy to the test. As others have reported, the strategy works very well for most of the illustrations: I zoomed through most every illustration the first book spending The trouble is when an outlier illustration comes along. When you're on an outlier illustration, you not only waste time following the path, but then you're left disoriented trying to trace back and worrying that you missed Waldo. Waldo-spotting performance degrades on these outlier illustrations, so don't be discouraged by Book 1, which has 4 (!) outliers in it. Overall, in terms of speed, this machine-learning based method is streets ahead of the old Slate strategy. Conclusions This was all done in good humor and -- barring a situation where someone puts a gun to your head and forces you to find Waldo faster than their colleague -- I don't recommend actually using this strategy for casual "Where's Waldo?" reading. As with so many things in life, the joy of finding Waldo is in the journey, not the destination. --- ## Python usage survey 2014 URL: https://www.randalolson.com/2015/01/30/python-usage-survey-2014/ Published: 2015-01-30 Categories: data visualization, python Tags: python, upgrade, version Randy Olson explores why Python users aren't upgrading to Python 3. Remember that Python usage survey that went around the interwebs late last year? Well, the results are finally out and I've visualized them below for your perusal. This survey has been running for two years now (2013-2014), so where we have data for both years, I've charted the results so we can see the changes in Python usage over time. I'll note that a big focus of this survey is to find out if Python users are transitioning over to Python 3, and if they aren't, then why they aren't making that transition. If you're on the fence about switching from Python 2 to 3, there are some great articles out there about the key differences between the two versions and the many, many, ...many advantages that Python 3 offers over Python 2. If you want to check out the underlying data yourself, head on over to Bruno Cauet's page. The first big question that we all we know is: How many Python users have actually used Python 2 and 3? As expected, nearly every Python user has used Python 2 at some point in their career. We also see good news for Python 3 over the past year: Python 3 usage increased 12 percentage points in 2014, up to nearly 3/4 of the surveyed Python users. Of course, "writing code with Python 3 once" doesn't mean that they actually use it regularly. The question below gets at that question more directly. Here we see yet more good news for Python 3: As much as 10% more Python users are primarily using Python 3 than in 2013, now accounting for 1/3 of all Python users. It seems the transition to Python 3 will be slow but steady for the next few years. The transition we see above may be caused by project managers at work. What version do Python users go to when working on their personal projects? Here we see a more close divide between the two versions: Nearly half of all users will start up their own projects with Python 3, whereas the other half still ardently remains pro-Python 2. Let's break these usage patterns down by the specific version now. Somehow, Python v. 2.5 and 2.6 are still in use in some places, but 2.7 still dominates the Python 2 landscape. We also see telltale signs that Python 3 is becoming a mature language in its own right, with users stuck in the older 3.2 and 3.3 versions. To summarize so far: Over 2/3 of all Python users are still using Python 2, with the majority of them sitting at 2.7. This tells us that Python is still very much a divided language, with a large portion of the user base unwilling to upgrade to the latest version despite the fact that Python 2 nearly 5 years old now. (That's ancient in programming language years!) A common complaint of the ardent Python 2 users is that Python 3 was a huge mistake. How does that turn out in this survey? Surprisingly, it seems the complainers are a minority: Only 12% of the Python users surveyed think Python 3 was a mistake. 1/3 of Python users think Python 3 is great (probably the ones who made the switch!), whereas over half of Python users think the Python developers could've made the Python 2 -> 3 transition more fluid. So, most Python users don't think Python 3 was a mistake, but 2 in 3 of them still haven't made the switch. That leaves us to wonder: Why haven't the Python 2 users made the switch yet? Package dependencies are -- by far -- the most-cited reasons for refusing to switch to Python 3. Many of us rely on specific packages -- such as numpy, scipy, pandas, etc. -- for our day-to-day work. If the packages haven't been upgraded to Python 3 yet, then why should we? Lucky for us, there's a web site dedicated to tracking the packages that are Python 3 ready. Are all of your packages on there? Then maybe it's time to make the switch. Another common reason for refusing to switch to Python 3 is that there's no incentive. We're comfortable with Python 2, we know its ins and outs, and everything works fine. Why bother upgrading? If you fall in this category, I'll point you to this article again that compares the key differences between Python 2 and 3. And don't forget about the 2 pounds of reasons why Python 3 is a huge upgrade over Python 2. The last two major reasons for refusing to switch to Python 3 are a little harder to address. If you have a large legacy code base or your manager simply refuses to make the upgrade, then it's up to you to convince management (or yourself) that the upgrade is worth the work. The links I included in the above paragraph are a good start. Finally, the survey wanted to look at cross-compatibility between Python 2 and 3. Here we start to see more good news for the future of Python 3: We saw a significant increase in the number of users who have ported their code from Python 2 to 3. It's not very well-known that there Python has packages for converting code from Python 2 to 3 and Python 3 to 2. When the survey asked users whether they'd used either of these packages even once, well, see for yourself... It seems like the Python devs need to do a better job of raising awareness of these code conversion packages. Because Python 2 and 3 really aren't that different, it's not necessary to write code for just one version. As the famous taco shell commercial goes: "Por que no los dos?" (Why not both?) To that end, the survey also polled users on whether they write Python 2- and 3- compatible code. Apparently only 1 in 3 Python developers bother to write multi-version compatible code, and there hasn't been much of a change since 2013. I look forward to seeing how the Python 3 transition progresses in 2014. If you have any more resources for helping (and convincing!) Python developers to make the transition, please share them in the comments below. --- ## The Shrinking Battleground: Every 4 years, fewer states determine the outcome of the Presidential election URL: https://www.randalolson.com/2015/01/12/the-shrinking-battleground-presidential-elections/ Published: 2015-01-12 Categories: data visualization Tags: electoral college, politics, presidential election, united states, voting Randy Olson delves into voting data to find the states that determine the outcome of the Presidential election. Every 4 years, Americans are tasked to elect the leader of one of the largest democratic nations in the world. Eager to have their voices heard, U.S. citizens from every state stand in line to cast their ballot for their favorite candidate. Yet, according to a research project spanning over a decade, many states have already selected their candidate long before a single ballot is cast. A shrinking battleground According to this project, an increasing number of states are consistently voting Democratic or Republican in every Presidential election. Below, I've broken down the states by how predictable they are over a period of 6 elections. A "Safe State" is a state that has a strong bias toward one political party (55% or more of the voting population favoring one party), whereas a "Swing State" has no such bias and could go either way. There has been a clear trend of declining swing states for the past 20 years: Nearly 2/3 of the states could've swung to either party in 1992, whereas only 14 of those states remained unpredictable in the 2012 Presidential election. Thanks to the winner-takes-all voting system in the U.S., if the majority of voters in a state vote for one candidate, that candidate receives all of the state's electoral votes regardless of how close the popular vote was. In essence, the winner-takes-all system all but guarantees 2/3 of the states' electoral votes to the political parties, so the only real competition in the Presidential election is in the few remaining swing states. Swing states and strongholds in the U.S. I was curious how the landscape of swing and safe states has looked over time, so I mapped the states below according to how many elections in a row they've voted for one party. This map handily shows the recent evolution of the American political landscape. The Mountain States and Alaska have long been a bastion of support for the Republican party, with the majority of them voting Republican for 11 or more Presidential elections in a row. The South is a more recent Republican acquisition, with most Southern states voting Republican in only the last 3-5 elections. On the other end of the political spectrum, the Northeast, Midwest, Western states, and Hawaii have remained strong supporters of the Democratic party for at least the past 6 elections. With most of the states (+ D.C.) basically carved out among the political parties, that leaves a measly 9 states that haven't consistently voted for one party for more than 2 Presidential elections. The tragedy of this situation is that -- in states where the outcome is all but decided -- the winner-takes-all system nullifies people's votes if they don't support the majority party in their state. Voting Democratic in Alaska is about as pointless as voting Republican in California, which is why so many voters don't bother showing up to the polls in these staunchly polarized states. Presidential campaign trails show us the states that matter What's worse is how even the Presidential candidates themselves indirectly acknowledge this fact during the Presidential elections. Over 50 years ago, the Presidential candidates visited most every state to gain nationwide support for their campaign. Yet in modern times, when we track both candidates' campaign trails during the 2012 Presidential election, it looks something like this: If you compare this map to the "Swing States and Strongholds" map above, you'll find that there's a near-perfect overlap. Ohio, Florida, Virginia, Iowa, and Colorado -- fairly well-known swing states nowadays -- receive a huge amount of attention from the Presidential candidates, whereas most of the other states are lucky to see the candidates in a political ad. FairVote's Presidential Tracker also shows the breakdown of ad spending in the 2012 election, which unsurprisingly falls in the same states. How can I make my vote count? If you live in a safe state, then it's in your best interest to see the winner-takes-all system done away with. As with all problems on a national scale, there's no easy solution. But FairVote seems to have a reasonable solution: Let's move to a National Popular Vote. --- ## A data-driven guide to creating successful reddit posts, redux URL: https://www.randalolson.com/2015/01/11/a-data-driven-guide-to-creating-successful-reddit-posts-redux/ Published: 2015-01-11 Categories: analysis, data visualization, reddit Tags: front page, reddit, upvotes Randy Olson revisits an old analysis to find out what makes for a successful reddit post. A couple years ago, I wrote an article using massive data set of reddit posts to tackle one of the more popular questions about reddit: How do I get a highly-upvoted post on reddit? In light of the recent findings that half of all reddit posts go ignored, this question is more important than ever. Given how much reddit has changed over the years, it seems appropriate to revisit this topic and see if my 2-year-old findings still hold true. Below, my undergraduate assistant Robert Bato and I analyzed 550,000 reddit posts from August 2014 to find out what makes a successful post on reddit nowadays. I bolded the big takeaway messages if you're not feeling like a long read. Disclaimer: I am only making statements about probability in this post. Following these guidelines will by no means 100% guarantee that you will have a successful post. Rather, following these guidelines will maximize your chance of having a successful post. When to post Just as last time, time is one of the most important factors when making a successful post on reddit. Below, we've charted the percentage of posts that receive a score higher than 100 -- a reasonable proxy for at least a modicum of success -- over the day of the week and time of day. We'll be referring to posts that receive a score higher than 100 as "successful posts" for the remainder of this article. As before, the U.S. Eastern time zone still rules reddit. Posts submitted around 9am EST tend to see the most success (around 5% garner >100 upvotes), whereas posts submitted around midnight EST typically do the worst (~2.5% garner >100 upvotes). I can only guess this is the result of American workers on the east coast coming into work and firing up reddit before actually starting their workday. Also note that the weekends see a larger fraction of successful posts, most likely because there's fewer posts submitted on the weekend (and therefore less competition). We've provided an alternative view of the same data below in a matrix format. Reading from side-to-side shows you the percentage of posts that scored >100 for a single day, whereas reading up-and-down shows you the percentage of posts that scored >100 for a particular time of day. The bluer the square is, the more posts that succeeded, whereas the whiter the square is the fewer posts that succeeded. The take-away for this section? To maximize your chances of having a successful post on reddit, you should submit your post around 9am EST on the weekends. If the weekends don't work for you, that's fine -- you can still achieve reasonable success by posting around 9am EST on any day. What to post Now that we know when to post, we need to figure out what to post. You may already have an article in mind that you want to share, in which case you can skip this section. But if you're looking for ideas on the kind of content to share, keep on reading. The most important thing to get out of the way is that imgur.com -- an image rehosting and sharing site -- is the domain of choice for most redditors. In fact, nearly 2/3 of all successful posts on reddit were links to an image on imgur. If you're looking to share an image, your best bet is to rehost it on imgur and share that link. If we rank all link domains submitted to reddit by the fraction of successful posts, we get the list below. Note that only link domains with at least 2,000 posts were included in this list so that it shows only the more popular domains on reddit. A quick scroll down the list shows reddit's love for sharing images and GIFs. If they're not rehosting images on imgur, they're sharing fancy HTML5 videos on gfycat, images from Wikimedia, or memes generated on LiveMeme. In fact, 14 of the top 25 domains shown here are dedicated to hosting and sharing some form of image or video. Wikipedia is especially a favorite on /r/TodayILearned, whereas ThinkProgress articles are a favorite to share all over reddit. Ars Technica is reddit's go-to place for all things tech, and no reddit gamer's life would be complete without a Steam account. reddit's favorite major news outlets also pop out on this chart: Time, Business Insider, and Slate. For better or for worse, your post stands the best chance of becoming successful if you're sharing some form of visual media: images, GIFs, or videos (in that order). Articles sometimes do well in certain subreddits depending on their focus, but they're almost always outperformed by pure image posts. I suspect this is because over half of reddit's users are on mobile devices, which makes reading lengthy articles unappealing. Where to post Now that we know to post some sort of visual media hosted on imgur around 9am EST, we need to figure out where to post it. For that, we looked at all subreddits that had at least 10,000 posts submitted to them (to make sure they're a decently active subreddit) and ranked them by the percentage of successful posts submitted there. /r/gonewild (NSFW) by far has the highest fraction of successful posts of the larger subreddits, but unless you're looking to share nude selfies, you probably don't want to post there. If you're looking to share some sort of image, /r/WTF (pics that make you go "WTF?"), /r/AdviceAnimals (silly memes), and /r/aww (cute animal pictures) are your top 3 choices. /r/TodayILearned is still the best place to share educational articles, and /r/Minecraft holds the bragging rights as a large gaming subreddit with the highest fraction of successful posts. Following the goal of this article, /r/WTF or /r/AdviceAnimals will give us the highest chance of having a successful image post. Don't bother posting in /r/gonewild or /r/soccer unless your image is relevant to those subreddits. Summary In sum, your post has the highest chance of success if you: Post it around 9am EST Post an image hosted on imgur Post an image on /r/WTF or /r/AdviceAnimals Of course, you don't need to follow these guidelines to the letter. Articles still do quite well on /r/TodayILearned and /r/worldnews, as do videos on (you guessed it) /r/videos. Give these tips a try for a week or two and report back how well they worked for you in the comments. Important caveat As always, it's vital that you should first pay close attention to the subreddit's posting rules before posting there. Posting a funny picture to /r/TodayILearned will only get your post removed by the moderators, so make sure your kind of post is allowed before submitting it. --- ## Over half of all reddit posts go completely ignored URL: https://www.randalolson.com/2015/01/11/over-half-of-all-reddit-posts-go-completely-ignored/ Published: 2015-01-11 Categories: data visualization, reddit Tags: front page, reddit, underprovision, voting Randy Olson delves into the data to find that over half of all reddit posts go completely ignored. A couple years ago, Eric Gilbert published a research article showing that more than half (~52%) of all popular links submitted to /r/pics go completely ignored the first time they're posted. I found this phenomenon to be strange because those 52% of links later went on to become wildly popular the second or third time they were posted, meaning that it's not just bad or uninteresting links that are going ignored on reddit: On a daily basis, we're missing out on hundreds of interesting links because no one bothered to upvote them. One drawback of Gilbert's study is that he focused solely on /r/pics. While /r/pics is one of the most popular subreddits, it certainly isn't representative of reddit as a whole. For that reason, my undergraduate assistant Robert Bato and I decided to follow up on Gilbert's findings to see if all of reddit was following suit. To do that, we collected every single post submitted to reddit in August 2014 and filtered them into 3 categories based on their score: >1 (received at least one upvote), =1 (received no votes at all). Below is the breakdown of those posts. If we use "Whenever you post a link on reddit, there's a 50/50 chance that it will be ignored (if you post like the typical redditor). It seems Gilbert was right on the money when he claimed that reddit is suffering from widespread underprovision. What does this mean for reddit? Logically, if we consistently find that half of all reddit posts are going ignored, that leaves us with the question: Has reddit reached its content carrying capacity? Is it impossible for reddit to effectively handle more than 50% of its posts on a daily basis? When I tallied the ignored posts based on defaults vs. non-defaults, the default subreddits accounted for about 2/3 of all the ignored posts on reddit. This tells us that the defaults are likely the main source of ignored posts: Only so many posts can receive adequate attention in a subreddit every day -- even a popular one -- and the rest are simply left to wither and die without a single upvote in reddit's archives. On the other end of the spectrum, the above statistic also tells us that 1/3 of all ignored posts are due to smaller subreddits that don't receive enough attention from users. These subreddits have users post links to them on a regular basis, but no one bothers to read the posts -- or at least upvote them. Both issues highlight that reddit still suffers from an over-concentration of users in the default subreddits. This conclusion is nothing new: I was writing about the flaws of reddit's default subreddit system almost 2 years ago. Despite reddit's continuous efforts to drive users into smaller subreddits, most users still focus on the defaults. So, what's the solution? I see two possibilities: reddit users need to upvote more and post less. If there's too many posts not receiving enough upvotes, then this is the obvious solution. But reddit's karma system doesn't reward people who upvote, so it seems unlikely that this will work with the current incentive system. reddit needs to filter more posts into smaller subreddits. The fewer posts a subreddit receives, the more attention (and possibly upvotes) each individual post will receive. Thus, we can increase reddit's content carrying capacity by filtering the posts typically going into the defaults into smaller, more niche subreddits. The problem with solution #2, of course, is that the reddit front page disproportionately rewards users who make it there. So even if there's an incredibly small chance of your post receiving any attention in a default subreddit compared to the smaller subreddits, if your post does reach the front page, it pays off more than if you posted in a smaller subreddit and more than makes up for all of your failed posts. (Game theorists might be thinking of risk analysis at this point...) Thus -- 2 years later -- I still believe reddit must do away with the default subreddit system if it stands a chance of reaching its maximum potential. A historical perspective on ignored reddit posts After I shared this article, /u/minimaxir repeated this analysis on all of reddit's posts back to 2008. I've shared his chart below with his permission. Between 2008 and 2010 -- reddit's nascent years -- upwards of 75% of all reddit posts went ignored. Once reddit started becoming popular in 2010 and beyond, we see the ratio of ignored posts go down to an even 50/50, then oddly stagnate there. This suggests to me that ~50% of all posts is reddit's carrying capacity: Regardless of the number of users that join reddit, the increased number of upvotes from new users will be offset by the increased number of posts they submit. As a consequence, just waiting it out won't solve reddit's underprovision problem; we need to rethink reddit's default subreddit structure. You can find /u/minimaxir's data source here. What does this mean for you? We all know that it's tough to have a successful post on reddit, but now we're beginning to learn that less than half of all posts receive just a single upvote -- an extremely loose criterion. This means that even if you make the best post possible on reddit, there's still a high chance that it's going to fail right out of the gate. Of course, if you pay attention to the reddit front page, there are some redditors who seem particularly talented at making successful posts on reddit -- and that's because there are some minor things you can do to increase your post's chances of being noticed. If you want to learn more about what factors into a successful reddit post, read this guide. --- ## Does a bigger film production budget result in more ticket sales? URL: https://www.randalolson.com/2014/12/29/does-a-bigger-film-production-budget-result-in-more-ticket-sales/ Published: 2014-12-29 Categories: analysis, data visualization Tags: movie budget, movie ticket sales, movies Randy Olson explores whether a bigger film production budget results in more ticket sales. If you take a stroll down a list of the most expensive films of all time, you'll notice that most of the films are from the past 15 years. Every year, more and more money is being poured into producing larger and grander films, likely with the hope that a bigger film production budget will result in more movie ticket sales. For the chart below, I analyzed the production budget and ticket sales data of 11,706 films (courtesy of Box Office Mojo) to see whether that assumption held up. (Note: All dollar values have been adjusted to 2014 dollars.) It's fun to note some of the outlier cases in the above chart: Midget Zombie Takeover is the most unremarkable film on this chart (in terms of budget and ticket sales), with a measly budget of $2,043 and ticket sales totaling $10,653. Tarnation is the biggest underdog success story with ticket sales totaling $721,546 coming from a budget of only $277. Zyzzyx Road is the biggest failure with a paltry $35 in ticket sales despite a $2.36 million production budget, although the poor ticket sales were due to a deliberate effort by the producer. Amusing outliers aside, the above chart is likely misleading because the few outlier cases in the lower budget and ticket sales ranges are strongly affecting the linear regression. For that reason, I focused the analysis only on big-budget films with at least $1 million in both budget and ticket sales. Although there's a weak correlation between film production budget and ticket sales (R^2 = 0.32), it's fairly clear that just pouring money into a film's production budget to hire high-profile actors, add more CGI, etc. doesn't mean that the film will sell more tickets. In fact, if we compare the linear regression (blue line) to the line of parity (black line), the more that's spent on film production, the less likely the film will end up making that investment back in ticket sales. Regardless, it's interesting to see that there's even a weak correlation between a film's budget and it's performance in the box office. Only 4 films with a $100M budget ever made less than $10M in ticket sales, whereas only 6 films with a budget less than $10M ever made more than $100M in ticket sales. It seems that the film industry is like the stock market: You have to spend money to make money. And if you want to make the real big money, you're gambling with an awful lot of benjamins. --- ## The biggest box office booms and busts since 1982 URL: https://www.randalolson.com/2014/12/29/the-biggest-box-office-booms-and-busts-since-1982/ Published: 2014-12-29 Categories: data visualization Tags: box office, movie ticket sales, movies Randy Olson charts out the films that have made and lost producers the most money over the past three decades. If you read my last post about the correlation between a film's budget and its performance in the box office, you were possibly intrigued about my mentions of the biggest box office successes and failures. I decided to focus on this topic a little more by using Box Office Mojo's data to provide some top 25 lists. For maximum alliteration, I'll call the successes "booms" and the failures "busts." The booms "Booms" are films that made the most money from ticket sales* after the cost of the production budget is subtracted. These are all films that you most likely went to see in the theaters at least once, if you were old enough. Interestingly, most of the films in this list are from the 1980s and 1990s. It seems that even though film production companies are spending more on producing bigger films, their investments aren't being matched by moviegoers at the theater. In fact, the only film from the past decade to make the top 25 is The Hunger Games, which (unsurprisingly) has had a sequel every year since it was released. Of course, the downside of looking at "the booms" by looking at net profit is that it favors the big-budget films with a massive marketing budget. What about the successful underdog films that were made in someone's bedroom with a low-grade camera? To find the underdog success stories, I calculated the profit ratio (net profit / budget) and ranked the films again. Paranormal Activity is by far the biggest underdog success story, having been shot on a $15,000 budget with a home video camera in a single house. The Blair Witch Project -- shot in a very similar manner to Paranormal Activity -- unsurprisingly shows up in 3rd place. Tarnation holds the record of the highest-profit film that was produced with less than $250. Incredibly, E.T. still shows up in the top 25 on this list despite its $10.5 million budget. Talk about a box office success! The busts "Busts" are films that had millions of dollars poured into them to hire high-profile actors, shoot stunning scenery, and produce the best CGI the film industry has to offer, but no one showed up in the theaters. You probably heard about these movies when they came out, then quickly forgot about them a few days later. Given the growing production budgets of modern films, it's no surprise that films from the past decade dominate this list. The biggest surprise in my mind is Waterworld, which suffered from an extremely bloated budget. Thankfully, Waterworld did much better in the international theaters and eventually broke even, but it's unlikely we'll see another Waterworld anytime soon. Another shocker on this list is Tangled, which ranks in as the most expensive animated film ever made with estimated production budget of $260 million. Again, Tangled eventually turned a profit when it was released in international theaters, but it's mind-boggling how expensive animated films can be! As with "the booms," focusing only on net losses limits the busts list to films with gargantuan production budgets that didn't live up to the producers' expectations. But what about the films that failed so spectacularly in the box office that they probably marked the end of the producers' career? To find the spectacular failures, I computed the loss ratio (net losses / budget) and ranked the films again. Don't be surprised if you've never heard of most of these films: They're on this list for a reason. Zyzzyx Road -- with a $1.3M budget -- holds a special place on this list because of its incredibly large loss ratio and it only brought in $30 in box office ticket sales. This spectacular failure was engineered by the producer, however, because he wanted to focus on international distribution of the film. Folks who consider The Boondock Saints a cult classic may be shocked to see it show up on this list. The Boondock Saints was terribly received when it was first released in 2000, and only later turned a profit when it developed a cult following long after its failure in the box office. That just goes to show that even though box office performance is typically used to gauge a film's success, the box office isn't the be-all and end-all of film fame. * When computing film profits, I summed all of the domestic ticket sales then, following the standard rule of thumb, divided the profits by half to account for movie theaters keeping a share of the ticket sales, taxes, etc. --- ## Top 10 grossing film studios in the U.S. (1982-2014) URL: https://www.randalolson.com/2014/12/29/top-10-grossing-film-studios-in-the-u-s-1982-2014/ Published: 2014-12-29 Categories: data visualization Tags: movie studios, movie ticket sales, movies Randy Olson charts the top grossing films studios in the U.S. Around this time last year, I kicked off a series of movie analysis posts to wrap up the year. In keeping with tradition, I figured I'd do the same this year. This year, I'll be looking at box office sales in the U.S. with data from Box Office Mojo. To kick things off, I thought it'd be interesting to look at the film studios that have brought in the most money from ticket sales in the past 30 years. It's fairly common to see lists of top-grossing films around the web, but it's much more rare to see the studios that are making bank off of them. The two biggest moneymakers by far are Warner Brothers and Buena Vista. If Buena Vista doesn't sound familiar, maybe their more popular brand name will ring a bell: Disney! Whereas Warner Brothers has profited mostly off of successful franchises (The Dark Knight, Harry Potter, The Hobbit, and more), Buena Vista has mostly drawn its profits from a huge variety of animated films such as Frozen, Toy Story, and Up. The line between the "have" and "have not" studios is drawn fairly clearly just after Sony, clearly explaining why these top 6 studios are referred to as the "big 6." Columbia and TriStar were merged into Sony in the early 2000s, and New Line Cinema was merged into Warner Brothers in 2008. The last independent film studio on this list -- MGM -- filed for bankruptcy in 2010, which hints at how tough it is to operate as a studio independent from the "big 6" nowadays. Of course, ticket sales are only a fraction of the profits that these studios bring in. As any parent can attest, film studios have become very efficient at extracting money from us through other merchandise as well, such as action figures, video games, candies, and so much more. --- ## Top 25 most gender-neutral names in the U.S. URL: https://www.randalolson.com/2014/12/06/top-25-most-gender-neutral-names-in-the-u-s/ Published: 2014-12-06 Categories: data visualization Tags: baby names, gender neutral, usa Randy Olson charts the top 25 most gender-neutral names in the U.S. As a long-time fan of Saturday Night Live, I have fond memories of the Pat sketch where Pat's friends were always trying to figure out his/her gender through a series of hilarious indirect tests. Despite their every effort -- from asking Pat which bathroom he/she uses to asking about love interests -- Pat's friends could never figure it out. Put on some techno music and watch that GIF for a little bit. It totally works. Part of the Pat skit relied on the fact that "Pat" is a fairly gender-neutral name. It could be short for Patricia or Patrick, so it's tough to tell Pat's gender from his/her name alone. As I was working through the U.S. baby names data set, the Pat sketch got me wondering: What are some other gender-neutral names like Pat? Below, I calculated the most gender-neutral names using entropy, which gives me the highest value (i.e., most gender-neutral) when the name is evenly distributed between boys and girls. Disappointingly, Pat doesn't even come close to the top 25. If we take how America names their babies as any indication of what Pat's gender was, only 11% of all babies named Pat in the last 30 years were female. America's parents have voted: Pat is a boy's name! There's a few more shockers on this list: Who names their child Infant or Baby?! The most likely explanation is that these are placeholder names until the parents/guardians can agree on a name. Apparently Justice is blind when it comes to gender. I feel bad for the 7,000+ boys named Elisha. At least there aren't many boys named Sue any more: 92% of all babies named Sue were female. Want to explore some more baby names on your own? Head on over to the Baby Name Explorer. Can you think of any more names that should be on this list? List 'em in the comments below. --- ## What caused the upsurge of unique American baby names in the 1970s? URL: https://www.randalolson.com/2014/12/06/what-caused-the-upsurge-of-unique-american-baby-names-in-the-1970s/ Published: 2014-12-06 Categories: analysis, data visualization Tags: african-americans, baby names, black power, usa Randy Olson explores the role of the Black Power movement in how American babies were named in the 1970s. Last week, I was exploring the ever-popular U.S. baby names data set and noticed a peculiar trend: The number of unique baby names has continued to rise dramatically for the last ~130 years -- with the exception of the past few years, of course. This observation could of course be explained by an increasing number of babies being born as time goes on. To account for that possibility, I plotted the unique number of baby names per 100,000 babies born. The 1880s had the most uniquely named babies, whereas the following generations saw a marked decline in distinctive baby names until the 1960s. It's possible that WWI followed by the Great Depression followed by WWII caused Americans to care less about giving their baby a distinctive name and more about keeping food on the table until next week. Whatever was the cause of the Great Baby Name Depression, it resulted in Baby Boomers being the least uniquely named generation in U.S. history. These charts can lead into a dozen different stories, but today I want to focus on only one story: What caused the upsurge of unique American baby names in the 1970s? During my search to find an answer, I ran across an interesting Freakonomics piece exploring whether a baby's name can predict their long-term income. During the discussion of the piece, a possible answer suddenly revealed itself: the Black Power movement of the 1960s and 70s. The Black Power movement Back in the 1960s and 70s -- only 50 years ago -- African-Americans faced overwhelming discrimination. At the height of the African-American Civil Rights Movement, a growing sect of African-Americans became disillusioned with the non-violent movement led by MLK and others. These African-Americans decided instead to take a more militant stance on the civil rights movement. This is the twenty-seventh time I have been arrested and I ain't going to jail no more! The only way we gonna stop them white men from whuppin' us is to take over. What we gonna start sayin' now is Black Power! - Stokely Carmichael (June 16, 1966) Prior to this movement, most African-American names closely resembled names used by European-Americans. As a part of the Black Power movement, many African-Americans began to openly embrace and take pride in their heritage and began giving their children uniquely African-American names. During this time, we began to see the rise of Arabic/Islamic names in the U.S.: (Note: The above name frequencies could be confounded with the fact that immigration was also on the rise from the 1970s onward; these aren't exclusively names given to African-American babies.) African-Americans also began to create their own names, later termed "inventive names." Prefixes such as La/Le, Da/De, Ra/Re, or Ja/Je and suffixes such as -ique/iqua, -isha, and -aun/-awn became common during this time. African-Americans even began adopting vocabulary terms as names for their children, a trend which still survives to this day. The Black Power and pride movement left an unmistakable mark on the U.S. historical record. It's interesting to note, however, that many of the "uniquely African-American" names that were popular in the 1970s and 80s have fallen out of favor in modern times. Have "uniquely African-American" names gone out of style, or have African-Americans adopted a different naming scheme? What about after the 1980s? Do my findings explain the continued rise of distinctive baby names in the 1980s and beyond? Certainly not. Immigration -- especially from Hispanic countries -- has undoubtedly played a huge role in the baby name landscape. There have also been several other "unique name" movements within the U.S. -- for example the recent "-Ayden" movement -- that are clearly visible in the historical record. But those are stories for next time. Want to explore some more baby names on your own? Head on over to the Baby Name Explorer. --- ## The most-viewed YouTube videos URL: https://www.randalolson.com/2014/12/03/the-most-viewed-youtube-videos/ Published: 2014-12-03 Categories: data visualization Tags: rankings, view count, YouTube Randy Olson charts out the most-viewed YouTube videos and highlights Gangnam Style's unprecedented rise to over 2 billion views. Earlier this week, Google announced that Psy's insanely viral YouTube music video, Gangnam Style, officially broke the 2,147,483,647 view barrier. What's significant about that number, you ask? That's the largest number that can be encoded by a 32-bit integer number. The folks at Google never suspected a video would exceed 2 billion views until Gangnam Style came along, essentially meaning that Gangnam Style is single-handedly responsible for forcing the Google team to upgrade YouTube's view counter to a 64-bit integer encoding. Considering the last time I watched Gangnam Style it was still in the measly 200 million view range, I was pretty shocked when I heard that it had recently broken 2 billion views. To put Gangnam Style's view count in perspective, I charted the top 30 most-watched YouTube videos of all time from Wikipedia: An estimated 3 billion people around the world -- roughly 40% of the world's population -- will be connected to the internet by the end of 2014. Gangnam Style has almost as many views as the number of people who have access to the internet, worldwide, and it's still growing. As a further testament to Psy's YouTube dominance, 2 of the top 10 -- and 3 of the top 30 -- most-viewed YouTube videos were made by Psy. The only artist to even come close to Psy is Katy Perry, but her three top 30 video view counts don't even add up to more than Gangnam Style's alone. Of course, the list above slightly favors older YouTube videos like Charlie bit my finger that were posted almost a decade ago. To give a better idea of the momentum that these videos have, I sorted the top 30 videos by the average number of views they received per day since the day they were posted. From this perspective, the view count race is far more competitive. Katy Perry's Dark Horse and Enrique Iglesias' Bailando -- both posted in early 2014 -- have managed to maintain the same "view count momentum" as Gangnam Style. Meanwhile, we see that oldies like "Charlie bit my finger" are really just resting on their laurels and will likely fall into YouTube obscurity by next year. As YouTube continues to grow, it will become increasingly common to see view counts in the billions. But Gangnam Style still holds the bragging rights for being the first to "break" YouTube. Update (December 7, 2014): Daniel Hadley did a great follow-up to this post looking at the most popular non-music videos on YouTube. --- ## Is reddit experiencing "upvote inflation"? URL: https://www.randalolson.com/2014/11/27/is-reddit-experiencing-upvote-inflation/ Published: 2014-11-27 Categories: data visualization Tags: front page, most upvoted, reddit, top posts, upvotes Randy Olson delves into 6 years of reddit posting data to discern whether reddit is experiencing "upvote inflation." A couple months ago, one redditor asked a simple question: "Is [reddit] karma experiencing inflation?" Of course, it's silly to think about inflation applying to upvotes because you can't buy anything with them. Nonetheless, it's interesting to find out whether an upvote on reddit is worth more or less nowadays than it did back in ye olden days of reddit. If we think of upvotes as a currency, we first need to know how many are in circulation at any given point in time. After all, the more upvotes that are around, the less your individual upvotes count. Below, I've charted the total number of upvotes that reddit posts received every day between 2008-2013. (Note the log scale of the y-axis.) Shown above, reddit posts receive orders of magnitude more upvotes today than they did back in 2008. Whereas a paltry 80,000 upvotes were spread amongst posts every day in early 2008, nowadays the reddit servers process well over 8 million upvotes every day. To put that in perspective: On average, reddit users pressed the upvote button 96 times every second in 2013. Essentially, this is what has been happening to the reddit servers: Needless to say, reddit upvotes have experienced massive inflation since 2008: An upvote nowadays is worth 113x less than it did in early 2008. That means you need 113x more upvotes than in 2008 to even stand a chance of reaching the front page. That also means that the karma hounds of the early reddit days, such as /u/qgyh2, have the real bragging rights of amassing huge mounds of upvotes when it was really hard to get them. Normally at this point I'd start making predictions about how much an upvote will be worth in the future, but there's actually been a strange trend starting in 2013: The total number of upvotes on reddit stagnated at about 8 million upvotes per day for all of 2013. Could this mean that reddit's upvote carrying capacity is about 8 million upvotes per day? If so, that's bad news for the multitude of posters than want their posts to be seen: As more and more users post links to reddit, an ever-growing percentage of them will go completely ignored. It should be telling to revisit this post next year to see if that trend held in 2014. Most upvoted posts Earlier this year, I wrote an article looking at the most upvoted posts of all time on reddit. Since it was much harder to amass upvotes in 2008 than today, it's necessary adjust every post's upvote count for the "upvote inflation" we see above. Below, I've charted out the most-upvoted posts on reddit every day when their upvote counts are adjusted to December 2013 upvotes. Interestingly, the most upvoted posts prior to mid-2011 garnered far more upvotes than posts today (relative to the total number of upvotes being given at the time). Perhaps the most extreme example is "test post please ignore", which amassed an incredible 1.56 million inflation-adjusted upvotes in July 2009. What happened in mid-2011 that changed everything? This change in the upvote distribution happens to coincide with the time period that /r/reddit.com was closed down. /r/reddit.com was undoubtedly the most popular catch-all subreddit back in the day, and the top post on /r/reddit.com was typically #1 on the front page. Without /r/reddit.com to dominate the #1 front page spot, we started to see a more even distribution of upvotes amongst the front page posts. Finally, here's the top 10 most-upvoted posts when their upvotes are adjusted for inflation: [1,559,193 upvotes] test post please ignore [1,014,411] For every upvote, I will donate $0.25 to the Julie Amero Defense Fund. Come on Redditors, let's do this! [968,303] For every upvote, I'll donate 10 cents to Darfur. [866,817] Vote up if you're NOT getting an iPhone today. [786,036] I hate my job... [785,427] Upvote this if you think marijuana should be legal. [750,413] Obama wins the Presidency! [731,547] Today is my birthday. I only get one every four years. I have a Karma score of 1... you know what to do! [726,002] Vote up if you think going to war with Iran is a bad idea [711,023] Vote up if you've lost all respect for Bill and Hillary Clinton (and now you see why "upvote if" threads have been banned by mods more-or-less site-wide.) You can download the entire list in csv format here. What about 2014?! I can't include many posts from 2014 in this analysis because they haven't been archived yet, which means that they can still be voted on. I'll be sure to update this article once all of the posts from 2014 have been archived. --- ## MMORPG Popularity, 1998-2013 URL: https://www.randalolson.com/2014/11/12/mmorpg-popularity-1998-2013/ Published: 2014-11-12 Categories: data visualization Tags: gaming, mmorpg, popularity Randy Olson charts the popular MMORPGs from 1998 through 2013. Back in my teenage years, I was absolutely addicted to MMORPGs. As soon as I got home from school, I was plugged into whatever virtual world I was currently bent on conquering with my comrades-in-arms. Suffice to say that when I ran across a data set of MMORPG subscriptions over time, I was quite keen on taking a data-centric look at how MMORPGs have risen and waned in popularity over time. Below, I charted the percentage of total MMORPG subscriptions every month for each MMORPG between 1998 and 2013. You'll notice that some MMORPGs don't even show up, and that's either because they didn't have a normal subscription-based model (à la Guild Wars) or the games were just an insignificant blip in the history of MMORPG subscriptions. If a MMORPG is missing that doesn't fall under those categories, then these guys are to blame. (Though feel free to mention it here in the comments!) Make sure to click on the image for a hi-res version It's crazy to think that only 20 years ago a vanishingly small fraction of people on Earth had access to the Internet. That of course meant that commercial MMORPGs didn't start popping up in earnest until the mid-late 1990s, when there were finally enough people rocking a 56k modem to actually play the games. One of the first graphical MMORPGs to pop up was The Realm Online, an unforgiving turn-based RPG developed by Sierra. Even though it was short-lived, The Realm Online still holds the bragging rights for being one of the oldest graphical MMORPGs that people still play to this day (even if it's only a few dozen players). What killed The Realm Online? Of course, it was the Great Grandaddy of all MMORPGs: Ultima Online. Widely regarded the first major commercial MMORPG, Ultima Online set the standard for what a MMORPG should be: Thousands of people to playing together at once, massively improved graphics, an extensive PvP system, and GMs who interacted with the players (for better or for worse). Noting the success of Ultima Online, EverQuest and Asheron's Call followed quickly thereafter, forming the "big three" MMORPGs of North America. Fun fact: EverQuest is infamous for being the MMORPG with the most expansion packs, with 21 expansions released over its 15 years of development. If you're wondering why you've never heard of (or at least, never played) Lineage despite its overwhelming popularity in the early noughties, that's because it was mostly popular in South Korea. Lineage holds the crown as the first MMORPG to reach 1 million subscribers, and was the first MMORPG to truly dominate the MMO world until World of Warcraft appeared on the scene. The early noughties saw dozens of MMORPG releases as game companies tried to cash in on the MMO wave. Dark Age of Camelot, Final Fantasy XI, Star Wars Galaxies... all came and went with mild success until the Big Daddy of MMORPGs appeared on the scene. World of Warcraft was to MMORPGs in North America what Lineage was to MMORPGs in South Korea. Until 2005, MMORPGs were the domain of hardcore gamers who were willing to dedicate a disgusting number of hours grinding to reach the highest echelons of the game. World of Warcraft finally made MMORPGs accessible to the casual gamer who perhaps only had an hour to play every other day. And the casual gamer rewarded the World of Warcraft developers handsomely: World of Warcraft is by far the most successful MMORPG in North America, at one point reaching 12 million monthly subscribers. To put that in context: Each monthly subscription costs $15, so Blizzard was raking in over $180 million every month from subscriptions alone. The late noughties saw several highly hyped but ultimately unsuccessful MMORPGs. (I'm looking at you, Age of Conan and Warhammer Online!) But it also saw the rise of several more niche MMORPGs: Second Life for objective-free play, RuneScape for web browser-based games, and EVE Online for masochists with a penchant for space travel. Meanwhile, Aion became the new Lineage in Asia, marking the third major commercial success for South Korea-based NCSOFT (after Lineage and Lineage II). So, what does the future hold for MMORPGs? Total subscriptions to MMORPGs have been falling since 2011, but most of those are due to World of Warcraft. Star Wars: The Old Republic (SWTOR) showed us that new MMORPGs can still take hold in the MMORPG ecosystem despite falling subscriptions. And if SWTOR and World of Warcraft are any indication, free-to-play is likely the future of MMORPGs as players grow ever unwilling to pay monthly fees to play a game. Hopefully this has been as fun and nostalgic journey for you as it has been for me. Now, time to load up that old EverQuest CD... --- ## Do women on OkCupid follow the Standard Creepiness Rule? URL: https://www.randalolson.com/2014/11/08/do-women-on-okcupid-follow-the-standard-creepiness-rule/ Published: 2014-11-08 Categories: data visualization Tags: dating, okcupid, standard creepiness rule, xkcd Randy Olson checks if women on OkCupid follow the Standard Creepiness Rule when it comes to dating. It seems that there's an XKCD comic for every life situation that we run in to. One of my favorites, by far, is the comic titled "Dating pools." This comic highlighted the Standard Creepiness Rule, a.k.a. the "half-your-age-plus-seven" rule, which states that no person should date someone under (age / 2 + 7), otherwise they will look like a creeper. This seems arbitrary, but if you crunch your age into that equation, I'm willing to bet that you wouldn't even consider dating someone under that age. (I would never consider dating someone under 21!) If we plot the Standard Creepiness Rule out for women: (Note that you can easily just change the axis labels in the above chart and it works just as well for men.) It just so happens that Christian Rudder released his book Dataclysm last month, which features a chart showing us the age range that women search on OkCupid for when looking for men to date. One of my first thoughts when I saw this chart was: Do women on OkCupid follow the Standard Creepiness Rule? (In case you're wondering about whether men follow the rule, check this post out.) Sure enough, if we overlay Rudder’s OkCupid data over the first chart, we see that women follow the rule almost exactly. There are a few spots in their teens and low-20s where women actually seek older men than the Standard Creepiness Rule says they should, but that trend quickly ends by the time they're 21. It's interesting to note that in their younger years, women seek older men much more than younger men and refuse to date men more than a year younger than them. This trend ends around their mid-30s, when women suddenly become okay with dating men up to 5 years younger than them. The most interesting trend reveals itself when you compare this chart to the male OkCupid user's chart: Whereas women tend to seek older men (in their younger years, at least), men are doing the exact opposite and seeking younger women. This results in an unfortunate situation for women as they age: During their younger years, they're highly sought after. But as they grow older, men's tendency to seek younger women leads to an ever-shrinking dating pool for older women. Check out Rudder's article on the topic to learn more. --- ## What makes for a stable marriage? Part 2 URL: https://www.randalolson.com/2014/11/06/what-makes-for-a-stable-marriage-part-2/ Published: 2014-11-06 Categories: data visualization, review Tags: dating, divorce, marriage, relationships, usa, wedding Randy Olson reviews more of a research paper that outlines what makes for a stable marriage in the U.S. About a decade ago, the gossip on everyone's lips was that "1/2 of all marriages in the U.S. end in divorce." That factoid was later disproven, but it left a lasting impression on the eligible bachelors and bachelorettes of America. In an effort to not become a part of that statistic, I started doing a little research on what makes for a stable marriage in America. Last month, I ran across an interesting study on divorce titled 'A Diamond is Forever' and Other Fairy Tales: The Relationship between Wedding Expenses and Marriage Duration. The authors of this study polled thousands of recently married and divorced Americans (married 2008 or later) and asked them dozens of questions about their marriage: How long they were dating, how long they were engaged, etc. After running this data through a multivariate model, the authors were able to calculate the factors that best predicted whether a marriage would end in divorce. What struck me about this study is that it highlighted about a dozen predictors that correlate with stable or unstable marriages in the U.S. By popular demand, I've highlighted 3 more of the biggest factors below as a follow-up to Part 1. I highly recommend checking the study out yourself (linked above) to look at all of them. First, I'll orient you on how to read these graphs. The authors always chose one category as the "reference point." That means that all of the other categories are compared to that category. Below for example, "59% less likely" means that couples who had a child before their engagement were 59% less likely to ultimately end up divorced than couples who did not have a child. Having children with your spouse We all know someone who was on the verge of a breakup or divorce until they announced that they were having a baby with their spouse. According to this study, having a baby with your spouse can decrease your chances of divorce by as much as 76% compared to couples who do not have children. Of course, having children within wedlock -- another telltale sign of a well-planned marriage -- reduces your chances of divorce moreso than having children before you tie the knot. What's particularly interesting, though, is that even having children out of wedlock still reduces your long-term chances of divorce. It seems that shotgun weddings are more stable than we would expect them to be! Being the same age as your spouse Perhaps unsurprisingly, the larger the age gap between you and your partner, the more likely your marriage will end in divorce. Only being 1-5 years away from your partner is nothing to worry about, but if you're old enough to be your partner's parent, then your marriage might be in trouble. Hugh Hefner, anyone? Note: A previous version of this article showed a chart giving specific relative percent likelihoods of divorce occurring based on number of years married. The original authors of the study have pointed out that although there is a significant correlation between wider age gaps and increased divorce, it is not possible to determine the relative percent likelihood from their study. That is left to future research. Having the same education level as your spouse If you're a PhD marrying a high school dropout, your marriage may be shakier than a marriage between two college graduates. It's particularly interesting to note that the education difference matters more for women than men: Women are 50% more likely to end up divorced when there is an education difference versus men at only 32% more likely. Important: correlation != causation Of course, it's important for us to keep in mind that these are all correlations with marriage stability, and they could be telling us any number of things. For example, the "having kids with your spouse" correlation could go either way: Either people in stabler marriages are more likely to have kids in wedlock, or people in less stable (unhappy) marriages tend not to have kids. All of the explanations I wrote above are my own interpretations of the correlations, but keep an open mind when thinking about what could really be driving these correlations with marriage stability. --- ## Where Democrats and Republicans want their tax dollars spent URL: https://www.randalolson.com/2014/11/06/where-democrats-and-republicans-want-their-tax-dollars-spent/ Published: 2014-11-06 Categories: data visualization Tags: democrats, opinion, politics, poll, republicans Randy Olson charts out where Democrats and Republicans want to see their tax dollars spent according to a nationally representative poll. Last week, I revealed the shocking age divide in where Americans want their tax dollars spent. This week, I'd like to focus on where Democrats and Republicans agree and disagree on how American tax dollars should be spent. To recap: I've been working with the UT Energy Poll on their latest poll for the past month, and I've had the chance to preview the issues that Americans think are important when voting for their representatives. This nationally representative poll asks Americans the question, "Where is it most important for the U.S. government to spend your tax dollars?," and they're given 8 options: Education Energy Environment Health care Infrastructure development/maintenance Job creation Military and defense Social Security We then break the answers down by various categories such as gender, political affiliation, level of education, etc. to see where Americans differ -- and agree -- in opinion. Here's how the poll results look when broken down by political affiliation: The options are ordered by overall preference -- with "Job creation" sitting at #1 and "Infrastructure" dead last -- to show where Americans want to see their tax dollars spent the most overall. The common contentions between Democrats and Republicans immediately reveal themselves in the middle of our national priority list: Democrats favor spending on health care, whereas Republicans would rather see those tax dollars spent on the military. Republicans are also heavily against spending on the environment, which is an unfortunate political development within the past 20 years. Perhaps the most unfortunate trend revealed in this chart, though, is that no political party seems to think that energy, the environment, nor infrastructure are particularly important to spend tax dollars on. Meanwhile, Social Security sits in second place with both Republicans and Democrats agreeing that it's important to pour tax dollars into. With a climate crisis on the horizon and America's failing infrastructure, we can only hope it's just a matter of time until can Americans get their priorities straight. Here's a version with a more color blind friendly color scheme: --- ## The evolution of chess openings and why GIFs make for bad data visualizations URL: https://www.randalolson.com/2014/11/02/the-evolution-of-chess-openings-and-why-gifs-make-for-bad-data-visualizations/ Published: 2014-11-02 Categories: data visualization Tags: chess, data visualization, GIF, video Randy Olson revisits the "evolution of chess openings" visualization to illustrate why videos and GIFs make poor visualization tools. Earlier this year, I went on a week-long data analysis frenzy into a massive data set of chess tournament games. One of the better visualizations that came out of that post series was the evolution of openings over time set, where I looked at the popularity of various chess openings from 1850 through 2014. However, once I tried to visualize up through the 4th move, I found it too difficult to use area charts any more. Instead, I turned to a GIF (or, well, a video-GIF): In retrospect, this video visualization is worthless for communicating how chess openings have changed over time. By watching the video, the only bit of information the viewer can possibly hope to glean is that some openings have dropped in popularity, and the openings have become more diverse over time. Most data visualizations take considerable cognitive resources to comprehend. The critical failure of this visualization is that on top of the typical cognitive resources required to interpret it, it also requires the viewer to remember what the graph looked like 5, 10, or even 30 seconds ago to make any sort of meaningful comparison. This requirement will undoubtedly lead to cognitive overload for the viewer, which ultimately renders the visualization unusable. In short: To make a better data visualization, show all of the important data at once. Don't require the viewer to remember parts of the graph from several seconds ago when making comparisons because it simply won't work. To illustrate this concept, I remade the graph as a stacked area chart in d3.js. I had to use smoothing to make the chart look reasonably presentable. Click on the image for a larger, fully labeled version Interactive version Once you're done looking through this visualization, go back to the video. Do you see any trends that you never saw before in the video? Is it easier to compare time points in this area chart than with the video? Keep these insights in mind the next time you consider using a GIF or video to visualize a time series. You can find the entire chess analysis series here: Part 1: Elo ratings Part 2: Game lengths and outcomes Part 3: Chess openings Part 4: Moves, captures, and checkmates The key to Magnus Carlsen's success --- ## The reigns and deaths of the Roman Emperors URL: https://www.randalolson.com/2014/10/29/the-reigns-and-deaths-of-the-roman-emperors/ Published: 2014-10-29 Categories: data visualization Tags: emperor, history, roman empire Randy Olson visualizes the reigns and deaths of the Roman Emperors over a 400 year period. Ah, the Roman Empire. One of the greatest empires to conquer the known world of ancient Europe. At its height, every man desired to sit on the throne, making the Roman Emperor one of the more precarious roles in ancient history. We always hear about the greatest and worst Emperors -- Caesar Augustus and Caligula, for example -- but we so rarely hear about the other Emperors who lost their lives in service of -- and sometimes in detriment to -- the Empire. A few weeks ago, someone on /r/dataisbeautiful charted out the length of the reigns of each King and Queen of England and how they ultimately met their end. (Spoiler: The royal line didn't seem very healthy.) Then earlier this week, another user charted out the causes of death of the Roman Emperors. It seemed only natural to combine the two to show the reigns and deaths of the Roman Emperors. All of this data comes from Wikipedia, who already had the entire line of Emperors sorted out into a table. Click on the image for a larger version In the visualization above, we see the fairly sordid history of the Roman Empire. Well over half of the Emperors met some form of premature and violent end, with the average reign lasting only 8 years throughout the history of the Empire. Most of these violent ends are attributable to a particularly unstable period of the Empire known as the Crisis of the Third Century, where over 20 men (Maximinus I -> Carinus) -- mostly prominent Generals of the army -- ascended to the throne in a mere 50-year period. Nearly all of the Emperors during this period died violently, either by assassination or in battle. The Crisis only came to an end when Diocletian, a military commander of low birth, came into power. Diocletian realized that the Empire was too large for one man to rule, and thus split the rule of the Empire with 3 other men, forming a tetrarchy, or "rule of four." Thanks to Diocletian's guidance and reforms, the Empire managed to hold together for another century despite The Crisis nearly tearing it apart. Diocletian is the only Emperor to have voluntarily abdicated from the throne, although it must have been tragic for him to watch the Empire crumble despite his life's work rebuilding it. The Roman Empire as we know it ended with Theodosius I who, despite his efforts, could not secure the Empire for his sons. His sons quickly lost control of the throne and the Empire fell into turmoil again, eventually leading to the sack of Rome in 410 AD and the eventual dissolution of the Western Roman Empire. The Eastern Roman Empire -- later known as the Byzantine Empire -- survived and even thrived for another millennium until it was conquered in 1453 AD. Want to learn more about Roman history? I highly recommend this excellent podcast called The History of Rome. --- ## The age divide in where Americans want their tax dollars spent URL: https://www.randalolson.com/2014/10/28/the-age-divide-in-where-americans-want-their-tax-dollars-spent/ Published: 2014-10-28 Categories: data visualization Tags: midterm elections, poll, united states, voting Randy Olson visualizes the age divide in where Americans want to see their tax dollars spent. With the U.S. Midterm Elections coming up, it's time again to rally your friends and family to fulfill their civic duty at the voting booths on November 4th. Considering the possibility that the Senate will flip parties, this will be a particularly important election to vote in. Coincidentally, I've been working with the UT Energy Poll on their latest poll for the past month, and I've had the chance to preview the issues that Americans think are important in this upcoming election. This nationally representative poll asks Americans the question, "Where is it most important for the U.S. government to spend your tax dollars?," and they're given 8 options: Education Energy Environment Health care Infrastructure development/maintenance Job creation Military and defense Social Security We then break the answers down by various categories such as gender, political affiliation, level of education, etc. to see where Americans differ -- and agree -- in opinion. One of the most striking disparities in opinion revealed itself when we looked at the responses by age: Percentages are the % of respondents in that age bracket The divide in priorities between young and old Americans couldn't be clearer: Older Americans overwhelmingly want their tax dollars spent on Social Security, military, and defense, whereas younger millennials prefer to see their tax dollars invested in job creation and education. This data adds to the pile of data demonstrating the growing divide between millennials and older generations. It seems that the millennial vs. older generation conflict reaches far beyond the working your way through college debate: millennials are tired of war-mongering in foreign countries and want to see those tax dollars invested at home instead, whereas their parents and grandparents are content to maintain the status quo as long as their own retirement is taken care of. With the tendency for 65+ year olds to turn out to vote far more than younger Americans, this divide could spell serious trouble for millennials who are struggling to find a job and pay off their college debt. Add in the fact that nearly half of the Senate falls into the 65+ age category and it's really no wonder that millennials feel vastly underrepresented in politics. Want to change the status quo? Get out to vote on November 4th. --- ## The reddit world map URL: https://www.randalolson.com/2014/10/27/the-reddit-world-map/ Published: 2014-10-27 Categories: data visualization, reddit, research Tags: internet, map, reddit, social network Randy Olson presents the interactive reddit world map. We can all agree that online social networks dominate most people's day-to-day Internet lives. 70% of all U.S. adults aged 18-29 have a Facebook account, and a large portion of those people check their Facebook at least once a day. What's strange is that most people regard social networks as nothing more than a blob of status updates and links that occasionally has something interesting on it. These people rely on word of mouth or rudimentary search features to find interesting content on social networks, despite the fact that there's often small communities focused on exactly what they want to talk about. Last year, I set out to change all that. I wanted to connect people to these smaller communities. The toughest part about online social networks is navigating them. With millions of users and hundreds of thousands of communities, how can we possibly hope to find the right community? That's when I had the thought: We all use maps to navigate the real world. Why can't we use maps to navigate online social networks? Randall Munroe famously inked a high-level social network map back in 2007: but his map wasn't particularly useful for navigating individual networks. Thus, I set out on an expedition to map out the untamed lands of my favorite social network -- reddit.com -- like a modern-day Lewis and Clark. Below are the results. The reddit world map After several months of web scraping, data tinkering, and fidgeting with layouts, I settled on a methodology for mapping reddit. I decided that I would project reddit onto a 2D plane where every subreddit is represented by a dot. A subreddit would connect to another subreddit if many users posted or commented in both of the subreddits, and the subreddit would be colored red if it connected to many subreddits or blue if it were connected to only a few. Finally, subreddits that connected to each other would be placed closer to each other on the 2D plane than subreddits that didn't, which had the neat effect of creating "meta-communities" of subreddits. When I initially shared this map, people started charting out the meta-communities that I had inadvertently created. Here's one of my favorite maps: Sure enough, if you zoom around on the interactive version of the map, you'll see these meta-communities of reddit pop up before your very eyes: video games, sports, technology, movies, music, and of course a huge porn peninsula. To better highlight these meta-communities, I decided to color the subreddits by their cluster. Community structure Unbelievably, those very same meta-communities popped out in the clustered version of the map. We see the active gaming cluster: The hub of reddit techies: The isolated My Little Pony island: And of course the massive porn cluster: Again, I made this interactive version available online to explore and make use of. All of this points to a major discovery: By Liking, commenting on, sharing, hashtagging, RTing, and voting on content in these social networks, we are creating a hidden structure that shows what we're really interested in, and what we're on the social network to talk about. These maps bring this hidden structure to light. If you'd like to read more about how these maps are constructed, you can read the corresponding research article here. Where do we go from here? Ultimately, I'd love to see this mapping methodology applied to social networks beyond reddit. Much like reddit and its subreddits, Twitter organizes itself around hashtags, Facebook organizes itself around Pages, and Pinterest organizes itself around Pins. The only thing stopping us from mapping these social networks out is the availability of their data. Come, my fellow Lewis and Clarks. Let's map out the untamed lands of social media. --- ## The ebb and flow of movies redux URL: https://www.randalolson.com/2014/10/26/the-ebb-and-flow-of-movies-redux/ Published: 2014-10-26 Categories: analysis, data visualization Tags: movies, streamgraph, ticket sales Randy Olson attempts to revive and update the popular "ebb and flow of movie ticket sales" graphic. Six years ago, the NY Times published one of my favorite interactive graphics entitled, "The ebb and flow of movies." This brilliant graphic concisely visualized over two decades of box office sales with the now-popular streamgraph. The colors indicated how well each movie did -- from hit to flop -- and the height of the vertical area signified how "hot" each movie was the box office each week. As you scroll through the years, you'll see now-classics like Back to the Future and Jurassic Park pop up and fade away in the seemingly endless stream of time. One of the most unfortunate parts of this graphic is that it's terribly outdated at this point: The latest movies listed were popular back when I was still a sophomore in undergrad. As such, this post will be the first in a series of articles documenting my attempts to revitalize and update "The ebb and flow of movies" to 2014 (and soon, 2015). Visualizing 3 decades of box office sales At this point, I've gathered the weekly box office sales of the top movies for the past 3 decades (1982-2014), which fits comfortably into a ~8 MB file. However, it's been surprisingly difficult to find visualization software that can visualize 3 decades of data into a streamgraph without crashing. I've had the best luck so far has been with the RAW engine, which I used to plot 5 years of box office sales (2009-2014) below. Darker blue = more total sales for that movie Lighter blue = less total sales Vertical height indicates weekly sales You can click on all of the images below for zoomed-in versions. Obviously this graph isn't particularly useful for communicating information about specific movies, but it does show the seasonal ebb and flow of movie sales, which I look at in more detail below. Ultimately, we'll need an interactive version similar to the NY Times rather than this static graph. I'm open to suggestions for software. In the meantime, I've visualized zoomed-in versions of the past few years of box office sales below. The left side of each label starts on its corresponding area. Sorry for the overlapping names in some cases; there are so many movies that I had to place all of them automatically. 2014 2013 2012 Seasonality of box office sales Another phenomenon that the original "Ebb and flow" article pointed to was the seasonality of box office sales: Summer blockbusters and holiday hits make up the bulk of box office revenue each year Streamgraphs make for a pretty presentation of sales, but the seasonal trends are better communicated with a line chart. Below, I calculated the median weekly movie sales for each week over the 1982-2014 period. Sure enough, we see the clear mark of the blockbuster season from mid-June through August and the rush of families to the movie theater after Thanksgiving and Christmas. What's more surprising to me is the spike in mid-April. What's going on there? That's all for now! Hopefully we can find the software to pull this revival off. --- ## Is 2014 the most unpredictable year of college football? URL: https://www.randalolson.com/2014/10/22/is-2014-the-most-unpredictable-year-of-college-football/ Published: 2014-10-22 Categories: data visualization Tags: AP Poll, college football, rankings, unpredictable Randy Olson charts the Top 25 AP Poll college football rankings to provide a historical perspective on how the top teams have fared. Florida State and Alabama have been dethroned. Several crowd favorites are struggling to stay in the top 25 (bye bye, Tigers!). And now the upstart Mississippi State Bulldogs -- who weren't ranked at the start of the season and have never won a national championship -- are ranked #1 in the national polls. It's no wonder that many spectators are starting to call 2014 one of the most unpredictable years in college football. I decided to delve into the AP rankings history a bit more to provide some perspective on how the rankings have changed over the course of the season. Below, I plotted each team's ranking from Week 1 (preseason) to Week 9 (this week). For clarity, I left out a few teams who only appeared in the top 25 for a single week. I've color coded the teams by their overall rankings trend: green for up, red for down, and grey for stable. The trend we see is quite impressive: Only 6 of the 30 teams who have appeared in the top 25 have maintained relatively stable rankings. 10 teams have gone up considerably in rank since the beginning in the season, leaving an astounding 14 teams who have suffered several upset losses throughout the season. How does 2014 compare to 2007? 2007 is widely regarded as one of the most unpredictable years in college football. So how does 2014 compare? Below, I charted out the full year of AP poll rankings for the teams who were in the top 25 at some point in 2007. (Again, for clarity I left out a few teams who only appeared for a week.) We see similar trends in 2007: Multiple teams coming out of nowhere vying for the #1 position, crowd favorites failing to perform (*cough* Wolverines *cough*), and the previous year's national champion falling in the rankings. If we tally the trend lines, we again see the signs of an unpredictable season: 15/40 teams up in the rankings, 15/40 down, and 10/40 stable. To make a direct quantitative comparison, I calculated the variance in each team's rankings from weeks 1-9 in 2007 and 2014. (Note: I counted unranked weeks for teams as a "26" ranking.) While 2007 saw the rankings change far more (average variance: 24), 2014 isn't too far behind (average variance: 19). It seems that it's a bit too early to definitively say that 2014 is the most unpredictable year in college football, but there's still plenty of big games left to shake things up. --- ## Who's responsible for gas prices at the pump? URL: https://www.randalolson.com/2014/10/17/whos-responsible-for-gas-prices-at-the-pump/ Published: 2014-10-17 Categories: data visualization Tags: commodity speculators, correlation, gas, oil, prices Randy Olson digs into the data to find out who's responsible for gas prices at the pump. With gas prices hitting an all-time low for 2014, many of us are left to wonder: Who's responsible for gas prices at the pump? According to the UT Energy Poll, the largest national poll on energy issues, 7 in 10 Americans believe that oil and gas companies are one of the biggest deciders in the price of gas. Given that so many Americans seem convinced on this issue, I decided to delve into gas prices a little further to see if this widespread belief held up to scrutiny. If you're in the know about gas prices, you probably haven't been surprised to see the reports that the price of crude oil has also been steadily declining for the past few months. What we see here is a correlation: As crude oil prices go down, gas prices quickly follow. The big question is: Does this correlation hold up over several years? Conveniently, the EIA publishes weekly crude oil and gas price estimates right on their web site. The weekly oil price estimates come out on Friday, and the weekly gas prices come out the following Monday. The chart below shows us how well the oil prices on Friday predict gas prices on Monday. Each dot represents a Friday-Monday pairing, and the line shows the ideal pairing if oil prices predicted gas prices perfectly. (Note: I didn't adjust for inflation here because I'm not making comparisons between time points.) What we see here is a near-perfect correlation over 21 years of oil and gas prices. In other words, oil and gas companies seem to have little to do with the price of gas; they simply follow the standard set by international crude oil prices. That's not to say that oil and gas companies don't add their own expenses to the cost of gas (as they should), but they aren't the ones responsible for those sudden spikes in gas prices that we see every now and then. If oil and gas companies aren't responsible for gas prices, then who is? According to the experts, commodity speculators play a huge role in determining the international price of crude oil. They keep an eye on the current worldwide supply and demand for crude oil, terrorist threats and disturbances, changing international markets, and several other indicators to best estimate the total supply and demand for crude oil in the future. In turn, these estimates directly affect the price of crude oil -- and ultimately, how much you pay at the pump. But that's a theory left to investigate another day. --- ## What makes for a stable marriage? URL: https://www.randalolson.com/2014/10/10/what-makes-for-a-stable-marriage/ Published: 2014-10-10 Categories: data visualization, review Tags: dating, divorce, marriage, relationships, usa, wedding Randy Olson reviews a research paper that outlines what makes for a stable marriage in the U.S. About a decade ago, the gossip on everyone's lips was that "1/2 of all marriages in the U.S. end in divorce." That factoid was later disproven, but it left a lasting impression on the eligible bachelors and bachelorettes of America. In an effort to not become a part of that statistic, I started doing a little research on what makes for a stable marriage in America. Earlier today, I ran across an interesting study on divorce titled 'A Diamond is Forever' and Other Fairy Tales: The Relationship between Wedding Expenses and Marriage Duration. The authors of this study polled thousands of recently married and divorced Americans (married 2008 or later) and asked them dozens of questions about their marriage: How long they were dating, how long they were engaged, etc. After running this data through a multivariate model, the authors were able to calculate the factors that best predicted whether a marriage would end in divorce. What struck me about this study is that it basically laid out what makes for a stable marriage in the U.S. I've highlighted 7 of the biggest factors below. I highly recommend checking the study out yourself (linked above) to look at all of them. How long you were dating First, I'll orient you on how to read these graphs. The authors always chose one category as the "reference point." That means that all of the other categories are compared to that category. Below for example, "20% less likely" means that couples who dated 1-2 years before their engagement were 20% less likely to ultimately end up divorced than couples who dated less than a year before getting engaged. What we see above is that dating 3 or more years before getting engaged leads to a much more stable marriage. This finding probably comes as no surprise, but it should stand as a warning to those who are eager to get married right away. Don't jump into marriage before you really get to know someone. How much money you make One depressing finding was that wealthier couples are less likely to end up divorced. The correlation couldn't be clearer: The more money you and your partner make, the less likely you are to ultimately file for divorce. How often you go to church Perhaps another important -- but unsurprising -- finding was that couples who attend church regularly have much stabler marriages. In fact, couples who never go to church are 2x more likely to divorce than regular churchgoers. Your attitude toward your partner If your partner's looks or wealth are an important factor in whether you want to marry them, then I've got bad news for you: Your marriage is more likely to end up in divorce than if you couldn't care less about wealth and good looks. These findings even more stereotypical when we break the categories down by gender. Men are 1.5x more likely to end up divorced when they care more about their partner's looks, and women are 1.6x more likely to end up divorced when they care more about their partner's wealth. How many people attended the wedding If you're following the above guidelines, you've been dating your partner at least 3 years before getting engaged, making a combined $125k salary, go to church together regularly, and don't worry about your partner's wealth nor looks. The Big Day is coming up and you're set to be happily married for life, right? Wrong! Crazy enough, your wedding ceremony has a huge impact on the long-term stability of your marriage. Perhaps the biggest factor is how many people attend your wedding: Couples who elope are 12.5x more likely to end up divorced than couples who get married at a wedding with 200+ people. Clearly, this shows us that having a large group of family and friends who support the marriage is critically important to long-term marital stability. How much you spent on the wedding The last graph would have us think that if we want a long-lasting marriage, we better be prepared to burn a hole in our pocket paying for a huge wedding. Yet the findings below completely contradict that intuition: The more you spend on your wedding, the more likely you'll end up divorced. The particularly scary part here is that the average cost of a wedding in the U.S. is well over $30,000, which doesn't bode well for the future of American marriages. In the research paper, the authors suggest that the financial burden incurred by lavish, expensive weddings leads to financial stress for the couple, which ultimately tears the marriage apart. They found that women, in particular, are vulnerable to divorce after expensive marriages: women in couples who spent $20,000 or more on their wedding are 3.5x more likely to end up divorced than their counterparts who spent less than half that. In other words, Bridezilla = Divorcezilla. Don't let advertisers fool you into spending your life savings on your wedding. Whether you had a honeymoon Whatever you do after your marriage, don't skimp on the honeymoon! Important: correlation != causation Of course, it’s important for us to keep in mind that these are all correlations with marriage stability, and they could be telling us any number of things. For example, the "how much money you make" correlation could go either way: Either people in stabler marriages are more likely to have a higher income, or couples with a low income could be more likely to divorce. All of the explanations I wrote above are my own interpretations of the correlations, but keep an open mind when thinking about what could really be driving these correlations with marriage stability. --- ## The most upvoted post on reddit every day URL: https://www.randalolson.com/2014/09/28/the-most-upvoted-post-on-reddit-every-day/ Published: 2014-09-28 Categories: data visualization, reddit Tags: front page, most upvoted, reddit, top posts Randy Olson analyzes over half a decade of reddit posting data to discover the most upvoted posts on reddit. One of the most commonly asked questions about reddit is, "What is the most popular reddit post of all time?" It's easy enough to go to /r/all/top and see the posts with the highest score, but thanks to reddit's vote fuzzing mechanism, a post's score isn't a reliable measure for how much attention the post received once it reached the front page. What we really want to know is which post received the most upvotes. The plot below is the result of crunching 6 years of reddit post data from 2008-2013. For each day, I've plotted total number of upvotes the most popular post received. I annotated a few posts that stood out. Top 10 reddit posts through 2013 This post wouldn't be complete without a top 10 list of the most upvoted reddit posts. Here they are. [240,730 upvotes] I am Barack Obama, President of the United States -- AMA [155,739] The safe. Some people doubted our resolve, but I said it would be open by New Years. [148,554] I’m Bill Gates, co-chair of the Bill & Melinda Gates Foundation. AMA [126,994] Please don't upvote, how do I remove the Skyrim mod "Schlongs of Skyrim"? [120,326] Can't believe what I found at the yard sale! [115,355] I sent Tom Hanks a 1934 Smith Corona typewriter with a typed invitation to come on my podcast. This was his response. [110,545] Tried taking a picture of a sink draining, wound up with a picture of an eye instead. [107,587] Mom was worried about my trip to the Grand Canyon, I sent her this picture. [105,272] After searching FB for people with the same name as me, I'd replicate their profile pic, make it my own and send them a friend request. Here are the pics [103,687] Airline screwed up, a friend just posted this on Facebook. History of top reddit posts The first truly popular post on reddit was "test post please ignore" on July 18, 2009. Back in the day, /u/qgyh2 was a reddit power user who first broke the record of reaching 100,000 karma. It seems he was making a test post to see how images showed up in the reddit comment threads. True to redditor fashion, the entire community rallied to reject his request to ignore the post, and "test post please ignore" became the most upvoted post on reddit for the next 2 years with an incredible 26,750 upvotes. The infamous "I will drink a beer for every upvote I get" thread soared to the upvote charts on St. Patrick's Day 2010, when redditors seemed intent on upvoting /u/chuckieballs under the table. By the time the thread was archived, /u/chuckieballs owed the reddit community 18,574 empty cans of beer. If /u/chuckieballs drank 11 beers every day between that day and today, he still wouldn't be finished fulfilling his promise. In 2011, we really start to see reddit become a major source for breaking news. With the announcements of Osama bin Laden's and Steve Jobs' deaths topping the front page of reddit before most major news sources knew what was happening, it's no wonder that reddit is often considered a major source of breaking news. Similarly in 2012 and 2013, we see major AMAs popping up on reddit, including President Obama's AMA on August 29, 2012 and Bill Gates' AMA on February 11, 2013. This chart really puts Obama's AMA into perspective: The amount of traffic that Obama's AMA drove to reddit was unprecedented and hasn't even come close to being matched since. Despite its growing popularity, reddit has maintained its sense of community that makes it so unique. For example, reddit ended 2013 with a bang with the announcement that The Mystery Vault had finally been opened. The top comment on the announcement thread summarizes the findings best: "Merry Christmas, reddit. You get nothing." What about 2014?! I can't include many posts from 2014 in this analysis because they haven't been archived yet, which means that they can still be voted on. I'll be sure to update this article once all of the posts from 2014 have been archived. --- ## Do men on OkCupid follow the Standard Creepiness Rule? URL: https://www.randalolson.com/2014/09/15/do-men-on-okcupid-follow-the-standard-creepiness-rule/ Published: 2014-09-15 Categories: data visualization Tags: dating, okcupid, standard creepiness rule, xkcd Randy Olson checks if men on OkCupid follow the Standard Creepiness Rule when it comes to dating. It seems that there's an XKCD comic for every life situation that we run in to. Is there an XKCD comic for that yet? One of my favorites, by far, is the comic titled "Dating pools." This comic highlighted the Standard Creepiness Rule, a.k.a. the "half-your-age-plus-seven" rule, which states that no person should date someone under (age / 2 + 7), otherwise they will look like a creeper. This seems arbitrary, but if you crunch your age into that equation, I'm willing to bet that you wouldn't even consider dating someone under that age. (I would never consider dating someone under 21!) If we plot the Standard Creepiness Rule out for men: (Note that you can easily just change the axis labels in the above chart and it works just as well for women.) It just so happens that Christian Rudder released his book Dataclysm last week, which features a chart showing us the age range that men search on OkCupid for when looking for women to date. One of my first thoughts when I saw this chart was: Do men on OkCupid follow the Standard Creepiness Rule? (And now we see why the last panel of the XKCD comic above applies so well to me...) Sure enough, if we overlay Rudder's OkCupid data over the first chart, we see that men follow the rule almost exactly. There are a few spots in the mid-30's where men seem willing to dip ever so slightly past the safe zone of non-creepiness, but that trend quickly ends by their 40's. Another interesting trend is how men aren't even close to reaching the upper bound of the zone of non-creepiness. According to the Standard Creepiness Rule, it'd be perfectly fine for a 30-year-old man to date a 45-year-old woman, but apparently 30-year-old men are already struggling with the idea of dating a 37-year-old! --- ## Who are the climate change deniers? URL: https://www.randalolson.com/2014/09/13/who-are-the-climate-change-deniers/ Published: 2014-09-13 Categories: data visualization Tags: america, climate change, denial, global warming, usa Randy Olson analyzes the latest American opinion poll results to answer the question, "Who are the climate change deniers?"" It seems that every 6 months, we see the news light up with reporters quoting the latest public opinion poll results on global climate change. "Climate change denial is up 7 percentage points this year." "1 in 5 Americans now deny that climate change is occurring." I've always wondered: Wouldn't it be much more helpful to provide a historical perspective on the issue? To accomplish just that, I took data from the renowned UT Energy Poll and plotted it below. The result? Despite the continuous raging debate, the average American's acceptance of global climate change hasn't really budged in the past 2 years. Acceptance of climate change in America has hovered around 70% since at least March 2012. Why is that? 97% of climate scientists agree that global climate change is occurring. The news was abuzz about climate change and global warming earlier this year. Even President Obama made an official statement on the topic in June. So why do these Americans refuse to budge on the issue? In a provocative article earlier this year, Chris Mooney suggested a simple answer: "Conservatives don't deny climate science because they're ignorant. They deny it because of who they are." Could this be true? Do climate change deniers know all there is to know about climate change and global warming, but they still refuse to accept it because of their beliefs? To get at this question, I dug further into the UT Energy Poll's most recent poll results. Household income and educational attainment First, I wanted to know if a person's income or educational attainment had anything to do with whether they accept that climate change is occurring or not. Thankfully, one of the demographic variables the poll collects is each person's household income. Due to the strong correlation between educational attainment and income, the graph below doubles as a measure of education. Surprisingly, income and educational attainment have no effect whatsoever on whether someone accepts that climate change is occurring or not: The "Yes" answers float around 70% regardless of whether the poll respondent brings in $20,000 or $200,000 a year. This finding seems to hint that knowledge has very little to do with whether someone accepts climate change nowadays. Political affiliation If knowledge has little to do with climate change acceptance, what does? American conservatives are well-known to be climate change deniers, so what do the poll results look like when we break the answers down by political affiliation? That's not an error with my plotting software. 86% of Democrats accept climate change, whereas half of all Republicans are still in denial on the issue. While there are still some Democrats in denial about climate change, it's fair to say that the majority of climate change deniers today are Republican. Religiosity Oddly enough, climate change has also become a religious issue in the past decade. Could a person's religiosity affect whether they accept that global climate change is occurring? Sure enough, we see the same trend as with political affiliation: The more religious a person is, the more likely they are to deny climate change. Whereas 80% of atheists accept climate change, only 56% of all very religious Americans agree. It's fairly clear from these graphs that religious, Republican American conservatives are the majority of climate change deniers today. If income, education, and knowledge has little to do with climate change acceptance, then could it be that climate change acceptance has become a cultural rather than factual issue in America? Do conservative Americans deny climate change simply because it conflicts with their identity as a conservative? If that's the case, then throwing facts at climate change deniers isn't going to make them budge on the issue. As Prof. Dan Kahan writes: "Everyone has gotten the memo on what 'climate scientists believe,'" and Mooney explains: If Kahan is right, the implication is that we need to talk about climate science in a way that is entirely devoid of cultural meanings that will antagonize the right. So what can we do to convince that last 30% of Americans? First off, we should stop talking about "what scientists believe" and instead actually take into consideration who we're trying to communicate with. Then we need to figure out: How can we discuss climate change without alienating the average American conservative? --- ## Clash of Clans Troop Efficiency URL: https://www.randalolson.com/2014/09/09/clash-of-clans-troop-efficiency/ Published: 2014-09-09 Categories: data visualization Tags: attack strategy, Clash of Clans, efficiency, troops Randy Olson analyzes all of the troops in Clash of Clans and discovers which ones give you the biggest bang for your buck. For the past couple months, I've been filling my downtime by playing Clash of Clans, a MMO base builder game where you can plunder other player's bases. In hopes of improving my attack strategy, I've read several guides on how to conduct the best attacks in the game. One thing that's missing from these guides, though, is an analysis of how efficient the attacks are. It's one thing to completely annihilate your opponent, but it's another challenge altogether to annihilate your opponent while expending as few resources as possible. To start up the conversation on attack efficiency in Clash of Clans, I've created several visualizations below looking at the cost and housing efficiency of every troop in the game. All of the data in these visualizations comes from here, and I've provided a tabulated version of the data here. You can access the interactive versions by clicking on the image. Troop cost efficiency Unless you're extremely patient, you're probably going to be raiding other player's bases to steal their resources so you can build up your own base. That means that you'll want to do as much damage to your opponent's base while spending as few resources as possible. Below, I've plotted the DPS and HP that every troop provides for every Elixir it costs. Each point represents an upgrade level for the troop. I had to plot the Elixir and Dark Elixir troops separately because there isn't a clear conversion rate between Elixir and Dark Elixir. (note that the axes in the chart above are on logarithmic scale to help display the data better) For experienced players, the main result here is likely unsurprising: Barbarians, Archers, and Goblins give you the biggest bang for your buck. Don't go raiding for resources with Giants, Wizards, or (worst) Dragons or you'll quickly find yourself losing resources every time you go raiding. It's hard to compare Elixir cost efficiency to Dark Elixir cost efficiency, but it's fairly clear from the plot above that Minions are the most cost effective Dark Elixir troop in the game. Witches may seem abysmal on paper until you remember that their main role is summoning Skeletons to attack the enemy base. What I found particularly interesting in the interactive versions of these plots is that almost every troop in the game becomes less cost efficient as you upgrade them. Upgraded troops may be stronger (and more fashionable), but you don't get as much out of them per Elixir. Troop housing space efficiency Of course, Clash of Clans isn't always about raiding villages for resources. Especially during Clan Wars, you'll want to completely wipe out your opponent's base so your clan gets more victory stars. In this case, your limiting resource becomes housing space. Every base can only support so many troops, and some troops take up more space than others. Below, I've plotted the DPS and HP that every troop provides for every housing space it takes up. The Golem stands tall as the ultimate tank in the game, whereas Wizards dish out the most damage per housing space (with Goblins as an underrated runner up). Higher-level Barbarians and Balloons seem to provide the best of both worlds, providing both ample DPS and HP for every housing space they take up. Again in the interactive version, you'll note that -- unsurprisingly -- all troops provide more DPS and HP per housing space as they're upgraded. This is ultimately the reason you should always upgrade your troops, even if they provide less bang for your purple buck. You'll also notice that some upgrades provide far more of an advantage than others. The level 5 Wizard upgrade improves the Wizard's DPS per housing space from 32 to 43 (+34%), whereas the level 6 upgrade only improves it to 45 (+5%), which hardly seems worth the exorbitant level 6 upgrade cost. The Balloon's level 5 and 6 upgrades take it from an underwhelming 14.4 DPS/HS & 56 HP/HS all the way up to 32.4 DPS/HS (+125%) & 109 HP/HS (+95%), marking them as two of the most cost effective upgrades in the game. Caveats There are of course some caveats to the results I outlined above. Particularly, if you followed the housing space efficiency chart, you may be inclined to spam Giants at your enemy and avoid Healers and expect to wipe out your opponent's base. Or worse, if you followed the cost efficiency chart, you may be inclined to spam Goblins and completely avoid Wall Breakers for the most cost efficient attack. Keep in mind that every troop has a particular purpose. Healers are useless by themselves, but they have great synergy with high-HP troops like the Giant. Goblins may have high DPS, but they attack defenses last and will die in droves if Wall Breakers don't open up the base for them. Think of these charts as a stepping stone, not a Rosetta stone, for crafting your perfect Clash of Clans attack strategy. --- ## Where the U.S. gets its oil from URL: https://www.randalolson.com/2014/08/28/where-the-u-s-gets-its-oil-from/ Published: 2014-08-28 Categories: data visualization Tags: energy, gasoline, imports, middle east, oil, usa Randy Olson dispels some common misconceptions about where the U.S. gets its oil from. Despite the fact that late-year gasoline prices have risen to the second-highest in recent memory, a new report from the UT Energy Poll shows that most Americans have little clue where their gasoline even comes from. According to the poll, 3 out of 4 Americans think that the U.S. imports the majority of its oil from somewhere in the Middle East. Yet when all of the U.S.'s oil imports are stacked up, oil from the Middle East comprises less than quarter of U.S. oil imports. In fact, the majority of the U.S.'s imported oil comes from countries in North and South America. If we look up the top 10 exporters of oil to the U.S., we might be surprised to find that our friendly neighbors to the north are the ones working the hardest to keep our gasoline tanks full. (Or at least, 87% of us will be surprised by this!) You're reading the chart right: the U.S. imported 701 million barrels of oil from Canada in only 6 months. That's roughly 13.3 billion gallons of gasoline! If you're a little more in-the-know than most people on this topic, you'll notice that the imports listed above don't even come close to matching the U.S.'s insatiable appetite for oil. And that's where the most important fact in this article comes in: 60% of the oil that Americans use is produced right here in the U.S. In fact, only between 1997-2010 did we see oil imports rise above our own oil production. This trend began to reverse in 2005 and we're now on a stable path toward (mostly) oil independence. It's time we put an end to this myth that the U.S. gets most of its oil from the Middle East. If you're one of today's lucky 10,000, I'm calling on you to share these facts with your family and friends today so they will be better informed when voting on energy policies in the future. Summary 3/4 of Americans don't know where the U.S. gets its oil from The U.S. only imports 40% of the oil it uses; the other 60% is produced in the U.S. Half of the oil the U.S. imports is from North and South America; less than 1/4 of imported oil comes from the Middle East --- ## The best and worst times to have your case reviewed by a judge URL: https://www.randalolson.com/2014/08/24/the-best-and-worst-times-to-have-your-case-reviewed-by-a-judge/ Published: 2014-08-24 Categories: data visualization, review Tags: decision making, depletion effects, judgement, mental energy, parole Randy Olson discusses the best and worst times to have your case reviewed by a judge, based on research data. Recently, I've been working my way through Daniel Kahneman's fascinating book on human decision making, "Thinking, Fast and Slow." In the third chapter, Kahneman discusses how external factors can affect our tendency to fall back on easy "default" decisions instead of taking a few minutes to think the decision through. To drive his point home, he provides one disturbing example from a research article published in PNAS a few years ago. In this article, the researchers analyzed 8 experienced judge's decisions on parole requests as a function of time of day. The judges reviewed about 35 cases per day, spending about 6 minutes on each case. On average, the judges approved only 36% of the parole requests presented to them each day, so the chances of having a positive judgement on your case are already bleak. Now, we'd expect judges -- of all people -- to be the best at making impartial decisions. If no external factors were affecting their decisions, we'd expect to see them consistently approving about 36% of the parole requests throughout the day. Let's take a look at what the researchers found. [caption align="aligncenter" width="550"] Proportion of parole requests approved as a function of what order they were reviewed in. Each tick on the x-axis denotes every third case. Circles denote when the judges took a food break. Source[/caption] Shockingly, the judges appear to be much more inclined to approve a parole request when they've just come off a break. In contrast, they reject far more requests than usual the closer they get to break time -- and nearly 100% of the requests just before they take a break. This study provides a classic example of depletion effects in human judgement, a theory which suggests that we have a limited amount of mental energy to expend during a working period. The longer we work on mentally strenuous tasks, the more mental energy we expend, and eventually we'll run out and start falling back to these easy -- and often wrong -- default decisions. In the judges' case, once a sufficient number of cases had worn them down, they started defaulting to rejecting every case put in front of them until it was break time. That means that perfectly eligible prisoners had to spend even more time in prison because the judge hadn't eaten his mid-morning snack. Yet more proof that humans don't make decisions in a vacuum: even missing breakfast can alter how you approach the day. The take-away here? Try your best to be seen by judges first thing in the morning or just after lunch. Take regular snack breaks throughout your workday; the longer you work without a break, the worse you perform. --- ## The myth of the smarter Atheist URL: https://www.randalolson.com/2014/08/24/the-myth-of-the-smarter-atheist/ Published: 2014-08-24 Categories: data visualization, review Tags: atheism, intelligence, IQ, myth, religion Randy Olson explains why we should put the myth of the smarter Atheist to rest. Ever since I published my previous article on the average IQ of students by college major, I've received several requests to analyze the correlation between IQ and religiosity. Below, I've written up an analysis of the existing published literature on the topic. I hope this serves as a springboard for future conversation on the topic -- and hopefully puts some tired myths to rest. Country-level evidence One of the few peer reviewed scientific articles I could find on the topic had the straightforward name, "Average intelligence predicts atheism rates across 137 nations." This article analyzes IQ and religiosity data from "137 countries that represent 95% of the world population" and claims to show that nations with more intelligent citizens tend to have more Atheists. I plotted the data from that article below. (for the statistics nerds, the R^2 on a linear regression = 0.352) What we see is a fairly weak relationship between national religiosity and average national IQ. Once we get up to about 20% of the population being Atheist, the IQ of the population flatlines at around 100 from then on. Even worse, in the ~0% Atheist range, there's a wide range of national IQs from 64 to 100+ -- with a cluster of low-IQ nations that appear to be driving the "trend." If we focus on the lowest IQ nations in the above chart, we notice that several of them are poor nations in e.g. Africa. That led me (and others who have reviewed the topic) to wonder whether the wealth of a nation better predicts the average intelligence of its citizens. After all, the more wealthy the average citizen is, the more time they have to dedicate to intellectual pursuits. (R^2 on a linear regression = 0.449; income per capita data from GapMinder) Indeed, if we look at income per capita instead of religiosity, we already see a much better correlation with average IQ. The correlation between religiosity and IQ is too weak to suggest that religiosity predicts intelligence on the national level. Anyone who claims otherwise is grasping at straws. Individual-level evidence Looking at national data is fine and dandy, but what about individual-level data? In another controversial research paper on the topic, Satoshi Kanazawa claimed to explain "Why Liberals and Atheists Are More Intelligent." In this article, Kanazawa analyzed data from a longitudinal study following students from middle school through adulthood. Conveniently, the study includes information on how religious the students are and the results of a standardized IQ test. Kanazawa mashed those variables together and claimed to show that Atheists tend to be smarter. I created a less misleading version of Kanazawa's plot below; judge the data for yourself. This is where we have to think about effect size vs. statistical significance. The most religious adults had an average IQ of 97.14, whereas the atheist adults had an average IQ of 103.09. That may seem like a wide gap -- 6 whole IQ points -- until we remember that anyone in the IQ range of 90-109 is classified as having "average intelligence." Thinking about this in practical terms: Would you be able to tell the difference between someone with a 97 IQ and someone with a 103 IQ? It's highly unlikely. So really, all Kanazawa showed is that the average person has average intelligence regardless of how religious they are. I'll leave the discussion of why this guy's work was published in the first place for the comments. Conclusion The take-away message? To my knowledge, no amount of research has shown that Atheists are notably smarter than highly religious folks. It's time we put this myth of the smarter Atheist to rest. --- ## How to make beautiful data visualizations in Python with matplotlib URL: https://www.randalolson.com/2014/06/28/how-to-make-beautiful-data-visualizations-in-python-with-matplotlib/ Published: 2014-06-28 Categories: data visualization, ipython, python, tutorial Tags: graphic design, matplotlib, python, tutorial, visualization Randy Olson provides code examples and explanations for a handful of beautiful data visualizations. Want to learn more about data visualization with Python? Take a look at my Data Visualization Basics with Python video course on O'Reilly. It's been well over a year since I wrote my last tutorial, so I figure I'm overdue. This time, I'm going to focus on how you can make beautiful data visualizations in Python with matplotlib. There are already tons of tutorials on how to make basic plots in matplotlib. There's even a huge example plot gallery right on the matplotlib web site, so I'm not going to bother covering the basics here. However, one aspect that's missing in all of these tutorials and examples is how to make a nice-looking plot. Below, I'm going to outline the basics of effective graphic design and show you how it's done in matplotlib. I'll note that these tips aren't limited to matplotlib; they apply just as much in R/ggplot2, matlab, Excel, and any other graphing tool you use. Less is more The most important tip to learn here is that when it comes to plotting, less is more. Novice graphical designers often make the mistake of thinking that adding a cute semi-related picture to the background of a data visualization will make it more visually appealing. (Yes, that graphic was an official release from the CDC.) Or perhaps they'll fall prey to more subtle graphic design flaws, such as using an excess of chartjunk that their graphing tool includes by default. At the end of the day, data looks better naked. Spend more time stripping your data down than dressing it up. Darkhorse Analytics made an excellent GIF to explain the point: (click on the GIF for a gfycat version that allows you to move through it at your own pace) Antoine de Saint-Exupery put it best: Perfection is achieved not when there is nothing more to add, but when there is nothing left to take away. You'll see this in the spirit of all of my plots below. Color matters The default color scheme in matplotlib is pretty ugly. Die-hard matlab/matplotlib fans may stand by their color scheme to the end, but it's undeniable that Tableau's default color scheme is orders of magnitude better than matplotlib's. Use established default color schemes from software that is well-known for producing beautiful plots. Tableau has an excellent set of color schemes to use, ranging from grayscale to colored to color blind-friendly. Which brings me to my next point... Many graphic designers completely forget about color blindness, which affects over 5% of the viewers of their graphics. For example, a plot using red and green to differentiate two categories of data is going to be completely incomprehensible for anyone with red-green color blindness. Whenever possible, stick to using color blind-friendly color schemes, such as Tableau's "Color Blind 10." Required libraries You'll need the following Python libraries installed to run this code: matplotlib pandas The Anaconda Python distribution provides an easy double-click installer that includes all of the libraries you'll need. Blah, blah, blah... let's get to the code Now that we've covered the basics of graphic design, let's dive into the code. I'll explain the "what" and "why" of each line of code with inline comments. Line plots import matplotlib.pyplot as plt import pandas as pd # Read the data into a pandas DataFrame. gender_degree_data = pd.read_csv("http://www.randalolson.com/assets/2014/06/percent-bachelors-degrees-women-usa.csv") # These are the "Tableau 20" colors as RGB. tableau20 = [(31, 119, 180), (174, 199, 232), (255, 127, 14), (255, 187, 120), (44, 160, 44), (152, 223, 138), (214, 39, 40), (255, 152, 150), (148, 103, 189), (197, 176, 213), (140, 86, 75), (196, 156, 148), (227, 119, 194), (247, 182, 210), (127, 127, 127), (199, 199, 199), (188, 189, 34), (219, 219, 141), (23, 190, 207), (158, 218, 229)] # Scale the RGB values to the [0, 1] range, which is the format matplotlib accepts. for i in range(len(tableau20)): r, g, b = tableau20[i] tableau20[i] = (r / 255., g / 255., b / 255.) # You typically want your plot to be ~1.33x wider than tall. This plot is a rare # exception because of the number of lines being plotted on it. # Common sizes: (10, 7.5) and (12, 9) plt.figure(figsize=(12, 14)) # Remove the plot frame lines. They are unnecessary chartjunk. ax = plt.subplot(111) ax.spines["top"].set_visible(False) ax.spines["bottom"].set_visible(False) ax.spines["right"].set_visible(False) ax.spines["left"].set_visible(False) # Ensure that the axis ticks only show up on the bottom and left of the plot. # Ticks on the right and top of the plot are generally unnecessary chartjunk. ax.get_xaxis().tick_bottom() ax.get_yaxis().tick_left() # Limit the range of the plot to only where the data is. # Avoid unnecessary whitespace. plt.ylim(0, 90) plt.xlim(1968, 2014) # Make sure your axis ticks are large enough to be easily read. # You don't want your viewers squinting to read your plot. plt.yticks(range(0, 91, 10), [str(x) + "%" for x in range(0, 91, 10)], fontsize=14) plt.xticks(fontsize=14) # Provide tick lines across the plot to help your viewers trace along # the axis ticks. Make sure that the lines are light and small so they # don't obscure the primary data lines. for y in range(10, 91, 10): plt.plot(range(1968, 2012), [y] * len(range(1968, 2012)), "--", lw=0.5, color="black", alpha=0.3) # Remove the tick marks; they are unnecessary with the tick lines we just plotted. plt.tick_params(axis="both", which="both", bottom="off", top="off", labelbottom="on", left="off", right="off", labelleft="on") # Now that the plot is prepared, it's time to actually plot the data! # Note that I plotted the majors in order of the highest % in the final year. majors = ['Health Professions', 'Public Administration', 'Education', 'Psychology', 'Foreign Languages', 'English', 'Communications\nand Journalism', 'Art and Performance', 'Biology', 'Agriculture', 'Social Sciences and History', 'Business', 'Math and Statistics', 'Architecture', 'Physical Sciences', 'Computer Science', 'Engineering'] for rank, column in enumerate(majors): # Plot each line separately with its own color, using the Tableau 20 # color set in order. plt.plot(gender_degree_data.Year.values, gender_degree_data[column.replace("\n", " ")].values, lw=2.5, color=tableau20[rank]) # Add a text label to the right end of every line. Most of the code below # is adding specific offsets y position because some labels overlapped. y_pos = gender_degree_data[column.replace("\n", " ")].values[-1] - 0.5 if column == "Foreign Languages": y_pos += 0.5 elif column == "English": y_pos -= 0.5 elif column == "Communications\nand Journalism": y_pos += 0.75 elif column == "Art and Performance": y_pos -= 0.25 elif column == "Agriculture": y_pos += 1.25 elif column == "Social Sciences and History": y_pos += 0.25 elif column == "Business": y_pos -= 0.75 elif column == "Math and Statistics": y_pos += 0.75 elif column == "Architecture": y_pos -= 0.75 elif column == "Computer Science": y_pos += 0.75 elif column == "Engineering": y_pos -= 0.25 # Again, make sure that all labels are large enough to be easily read # by the viewer. plt.text(2011.5, y_pos, column, fontsize=14, color=tableau20[rank]) # matplotlib's title() call centers the title on the plot, but not the graph, # so I used the text() call to customize where the title goes. # Make the title big enough so it spans the entire plot, but don't make it # so big that it requires two lines to show. # Note that if the title is descriptive enough, it is unnecessary to include # axis labels; they are self-evident, in this plot's case. plt.text(1995, 93, "Percentage of Bachelor's degrees conferred to women in the U.S.A." ", by major (1970-2012)", fontsize=17, ha="center") # Always include your data source(s) and copyright notice! And for your # data sources, tell your viewers exactly where the data came from, # preferably with a direct link to the data. Just telling your viewers # that you used data from the "U.S. Census Bureau" is completely useless: # the U.S. Census Bureau provides all kinds of data, so how are your # viewers supposed to know which data set you used? plt.text(1966, -8, "Data source: nces.ed.gov/programs/digest/2013menu_tables.asp" "\nAuthor: Randy Olson (randalolson.com / @randal_olson)" "\nNote: Some majors are missing because the historical data " "is not available for them", fontsize=10) # Finally, save the figure as a PNG. # You can also save it as a PDF, JPEG, etc. # Just change the file extension in this call. # bbox_inches="tight" removes all the extra whitespace on the edges of your plot. plt.savefig("percent-bachelors-degrees-women-usa.png", bbox_inches="tight") Line plots with error bars import pandas as pd import matplotlib.pyplot as plt from scipy.stats import sem # This function takes an array of numbers and smoothes them out. # Smoothing is useful for making plots a little easier to read. def sliding_mean(data_array, window=5): data_array = array(data_array) new_list = [] for i in range(len(data_array)): indices = range(max(i - window + 1, 0), min(i + window + 1, len(data_array))) avg = 0 for j in indices: avg += data_array[j] avg /= float(len(indices)) new_list.append(avg) return array(new_list) # Due to an agreement with the ChessGames.com admin, I cannot make the data # for this plot publicly available. This function reads in and parses the # chess data set into a tabulated pandas DataFrame. chess_data = read_chess_data() # These variables are where we put the years (x-axis), means (y-axis), and error bar values. # We could just as easily replace the means with medians, # and standard errors (SEMs) with standard deviations (STDs). years = chess_data.groupby("Year").PlyCount.mean().keys() mean_PlyCount = sliding_mean(chess_data.groupby("Year").PlyCount.mean().values, window=10) sem_PlyCount = sliding_mean(chess_data.groupby("Year").PlyCount.apply(sem).mul(1.96).values, window=10) # You typically want your plot to be ~1.33x wider than tall. # Common sizes: (10, 7.5) and (12, 9) plt.figure(figsize=(12, 9)) # Remove the plot frame lines. They are unnecessary chartjunk. ax = plt.subplot(111) ax.spines["top"].set_visible(False) ax.spines["right"].set_visible(False) # Ensure that the axis ticks only show up on the bottom and left of the plot. # Ticks on the right and top of the plot are generally unnecessary chartjunk. ax.get_xaxis().tick_bottom() ax.get_yaxis().tick_left() # Limit the range of the plot to only where the data is. # Avoid unnecessary whitespace. plt.ylim(63, 85) # Make sure your axis ticks are large enough to be easily read. # You don't want your viewers squinting to read your plot. plt.xticks(range(1850, 2011, 20), fontsize=14) plt.yticks(range(65, 86, 5), fontsize=14) # Along the same vein, make sure your axis labels are large # enough to be easily read as well. Make them slightly larger # than your axis tick labels so they stand out. plt.ylabel("Ply per Game", fontsize=16) # Use matplotlib's fill_between() call to create error bars. # Use the dark blue "#3F5D7D" as a nice fill color. plt.fill_between(years, mean_PlyCount - sem_PlyCount, mean_PlyCount + sem_PlyCount, color="#3F5D7D") # Plot the means as a white line in between the error bars. # White stands out best against the dark blue. plt.plot(years, mean_PlyCount, color="white", lw=2) # Make the title big enough so it spans the entire plot, but don't make it # so big that it requires two lines to show. plt.title("Chess games are getting longer", fontsize=22) # Always include your data source(s) and copyright notice! And for your # data sources, tell your viewers exactly where the data came from, # preferably with a direct link to the data. Just telling your viewers # that you used data from the "U.S. Census Bureau" is completely useless: # the U.S. Census Bureau provides all kinds of data, so how are your # viewers supposed to know which data set you used? plt.xlabel("\nData source: www.ChessGames.com | " "Author: Randy Olson (randalolson.com / @randal_olson)", fontsize=10) # Finally, save the figure as a PNG. # You can also save it as a PDF, JPEG, etc. # Just change the file extension in this call. # bbox_inches="tight" removes all the extra whitespace on the edges of your plot. plt.savefig("chess-number-ply-over-time.png", bbox_inches="tight"); Histograms import pandas as pd import matplotlib.pyplot as plt # Due to an agreement with the ChessGames.com admin, I cannot make the data # for this plot publicly available. This function reads in and parses the # chess data set into a tabulated pandas DataFrame. chess_data = read_chess_data() # You typically want your plot to be ~1.33x wider than tall. # Common sizes: (10, 7.5) and (12, 9) plt.figure(figsize=(12, 9)) # Remove the plot frame lines. They are unnecessary chartjunk. ax = plt.subplot(111) ax.spines["top"].set_visible(False) ax.spines["right"].set_visible(False) # Ensure that the axis ticks only show up on the bottom and left of the plot. # Ticks on the right and top of the plot are generally unnecessary chartjunk. ax.get_xaxis().tick_bottom() ax.get_yaxis().tick_left() # Make sure your axis ticks are large enough to be easily read. # You don't want your viewers squinting to read your plot. plt.xticks(fontsize=14) plt.yticks(range(5000, 30001, 5000), fontsize=14) # Along the same vein, make sure your axis labels are large # enough to be easily read as well. Make them slightly larger # than your axis tick labels so they stand out. plt.xlabel("Elo Rating", fontsize=16) plt.ylabel("Count", fontsize=16) # Plot the histogram. Note that all I'm passing here is a list of numbers. # matplotlib automatically counts and bins the frequencies for us. # "#3F5D7D" is the nice dark blue color. # Make sure the data is sorted into enough bins so you can see the distribution. plt.hist(list(chess_data.WhiteElo.values) + list(chess_data.BlackElo.values), color="#3F5D7D", bins=100) # Always include your data source(s) and copyright notice! And for your # data sources, tell your viewers exactly where the data came from, # preferably with a direct link to the data. Just telling your viewers # that you used data from the "U.S. Census Bureau" is completely useless: # the U.S. Census Bureau provides all kinds of data, so how are your # viewers supposed to know which data set you used? plt.text(1300, -5000, "Data source: www.ChessGames.com | " "Author: Randy Olson (randalolson.com / @randal_olson)", fontsize=10) # Finally, save the figure as a PNG. # You can also save it as a PDF, JPEG, etc. # Just change the file extension in this call. # bbox_inches="tight" removes all the extra whitespace on the edges of your plot. plt.savefig("chess-elo-rating-distribution.png", bbox_inches="tight"); Easy interactives As an added bonus, thanks to plot.ly, it only takes one more line of code to turn your matplotlib plot into an interactive. More Python plotting libraries In this tutorial, I focused on making data visualizations with only Python's basic matplotlib library. If you don't feel like tweaking the plots yourself and want the library to produce better-looking plots on its own, check out the following libraries. Seaborn for statistical charts ggplot2 for Python prettyplotlib Bokeh for interactive charts Recommended reading Edward Tufte has been a pioneer of the "simple, effective plots" approach. Most of the graphic design of my visualizations has been inspired by reading his books. The Visual Display of Quantitative Information is a classic book filled with plenty of graphical examples that everyone who wants to create beautiful data visualizations should read. Envisioning Information is an excellent follow-up to the first book, again with a plethora of beautiful graphical examples. There are plenty of other books out there about beautiful graphical design, but the two books above are the ones I found the most educational and inspiring. If you liked what you saw in this post and want to learn more, check out my Python data visualization video course that I made in collaboration with O'Reilly. In just one hour, I will cover these topics and much more, which will provide you with a strong starting point for your career in data visualization. --- ## Average IQ of students by college major and gender ratio URL: https://www.randalolson.com/2014/06/25/average-iq-of-students-by-college-major-and-gender-ratio/ Published: 2014-06-25 Categories: data visualization Tags: college major, gender differences, IQ score, SAT, standardized tests Randy Olson charts out the IQ of students by their major's gender ratio and reveals a startling trend. After all the controversy that arose after I posted my breakdown of college majors by gender last week, I promised myself I'd stay away from controversial gender-related topics for a while. But when I ran across an ETS-curated data set of average student IQs by college major, I couldn't avoid putting this visualization together. Below, I plotted several college major's estimated average student IQ over the gender ratio of that major. The result? A shockingly clear correlation: the more female-dominated a college major is, the lower the average IQ of the students studying in the major. A naive reader may look at this graph and conclude that men are smarter than women, but it is vital to note that, on average, men and women have about the same IQ. By popular request, here's an interactive version of the above chart: https://plot.ly/~etpinard/330/us-college-majors-average-iq-of-students-by-gender-ratio/ IQs are typically classified as follows: 130+: Very superior intelligence 120-129: Superior 110-119: Above average 90-109: Average Considering that many of the female-dominated majors heavily involve interpersonal interactions, my initial thought was that this all made sense: Women are widely known to be more socially-inclined and nurturing than men, so we would expect to see them dominate fields that heavily involve people. But how does that explain the drastic IQ differences between male- and female-dominated fields, if the average man and woman have the same IQ? The answer comes from the fact that the IQ score here is estimated from the students' SAT score. This isn't an altogether unreasonable approach: Several studies have shown a strong correlation between SAT scores and IQ scores. But if we break down the SAT score by Verbal and Quantitative, we see why this IQ estimation is potentially misleading. If we re-make the first plot against the Verbal SAT score, we see that it's basically a wash: there's no correlation between a major's gender ratio and the average student's Verbal SAT score. When we plot the students' Quantitative SAT score against the major's gender ratio, we see the negative correlation appear again. This tells us that the original plot is actually showing preference for quantitative majors: The higher the estimated IQ, the more quantitative/analytical the major, and the fewer women enrolling in those majors. This brings up an interesting question of how valuable the SAT is as a standardized test across all majors, if a higher SAT score is really only indicating that the student is better at solving quantitative/analytical problems. Not all majors require a high analytical aptitude, after all. Technical bits Some of my readers requested the R^2 for the above plots. Here they are: The R^2 on the IQ vs major's gender ratio graph is 0.601 The R^2 on the Verbal SAT vs. major's gender ratio graph is 0.019 The R^2 on the Quantitative SAT vs. major's gender ratio graph is 0.738 The R^2 between Quantitative SAT score and Verbal SAT score is 0.027 For those who want to know what R^2 means: http://en.wikipedia.org/wiki/Coefficient_of_determination Notice about the IQ data Since I posted this article, the veracity of the IQ data set has been brought into question. I think StatisticBrain is a fairly reliable data source, but I write this here so readers can come to their own opinion about what this data shows, and how much to trust it. The data source says "Graduate Record Examination scores" then goes on to list SAT scores. Which is it? According to this comment, the scores listed are pre-2011 GRE scores, which can be found on the ETS web site here. The IQ estimates appear to have been performed separately from ETS, perhaps by StatisticBrain. So what does this mean for the graphs above? The IQ estimates are representative of students who are in their last year of undergraduate studies (or have already graduated) and are intending to apply to one of the majors. That makes the IQ estimates an imperfect sample, as some students may be changing majors for graduate school. I'd like to see this analysis redone with the SAT scores of students tied to their final undergraduate college major rather than intended graduate school major. --- ## Why the Dutch are so tall URL: https://www.randalolson.com/2014/06/23/why-the-dutch-are-so-tall/ Published: 2014-06-23 Categories: data visualization Tags: diet, dutch, historical height, nutrition, the netherlands Randy Olson charts out the median male height in various countries from 1820-2013 and explains why the Dutch are so tall. It's fairly common knowledge that the Dutch are some of the tallest people in the world. Whereas the average American man measures in at about 5'9" (176 cm), the average Dutch man stands at well over 6' (185 cm) tall. What is it about this small, traditionally seafaring nation that breeds such extraordinarily tall people? Contrary to popular belief, it's not to keep their heads above water. To provide a historical perspective, I charted the median male height for various countries between 1820 and 2013 below. It was surprisingly difficult to find this kind of height data, but fortunately many of these country's militaries meticulously recorded the median height of their new conscripts every year. These records provide a convenient (albeit somewhat biased) sample of the young generation of men during the time period. The raw data for this chart is available on figshare here. You'll notice that there's several holes in the data set, which I simply extrapolated the trends over. I compiled this data set from half a dozen different sources, so if you plan to use this data set for any of your research, I strongly suggest double-checking the sources I list there. The most surprising revelation here is that the Dutch became the tallest Europeans only recently in the 1980s. Before then, they were one of the shortest people in Europe at only 5'5" (165 cm) for first half of the 19th century. What changed after 1850 that led to this explosive Dutch growth? Prof. Drukker at the University of Groningen suggests that it has a lot to do with the distribution of wealth. As Cecily Layzell writes: The Dutch growth spurt of the mid-19th century coincided with the establishment of the first liberal democracy. Before this time, [The Netherlands] had grown rich off its colonies but the wealth had stayed in the hands of the elite. After this time, the wealth began to trickle down to all levels of society, the average income went up and so did the height. This explanation makes intuitive sense: It's well-known that we're much taller than our ancestors 100 years ago because of improved nutrition, especially in our adolescent years. If the average citizen has more money to buy healthy food, then we would expect their children to grow bigger, stronger, and taller. To add more evidence to the pile: GapMinder clearly shows that the Dutch income per capita stagnated until the mid-late 19th century, right when the Dutch median height started rising as well. So, there we have it. Make sure all of our citizens are wealthy enough to buy healthy food and their children will grow up to be bigger, stronger, and healthier. It's not as fun an answer as we would've hoped for, but at least we can put this "head above water" theory to rest! Edit (6/29/2014): Several of my readers have rightly pointed out that although this data explains why many Europeans have grown taller in the past 150 years, it doesn't necessarily explain why the Dutch are so much taller than the rest of Europe. There are a couple possibilities that merit further investigation: The Dutch diet: The average Dutch citizen eats a lot of breads, meats, cheese, and drinks a lot of milk -- moreso than many of their European counterparts. The Dutch genes: It's fairly well-known that pre-civilization humans were much taller than their civilized counterparts. It's possible that the Dutch ancestors from thousands of years ago were always taller, but Dutch diet and nutrition limited how large they grew. That still leaves open the question of why the Dutch ancestors were taller than the rest, however. --- ## Accuracy of three major weather forecasting services URL: https://www.randalolson.com/2014/06/21/accuracy-of-three-major-weather-forecasting-services/ Published: 2014-06-21 Categories: data visualization Tags: accuracy, meteorologist, nate silver, prediction, rainfall, weather forecast Randy Olson charts out the accuracy of three major weather forecasting services. For the past month, I've been slowly working my way through Nate Silver's book, The Signal and the Noise. It's really a great read, but if you're a regular reader on this blog, I'd imagine you've already read it. This book is loaded with all kinds of great examples of where predictive analytics succeeds and fails, and I decided to highlight his weather forecasting example because of how surprising it was to me. For those who aren't in the know: Most of the weather forecasts out there for the U.S. are originally based on data from the U.S. National Weather Service, a government-run agency tasked with measuring and predicting everything related to weather across all of North America. Commercial companies like The Weather Channel then build off of those data and forecasts and try to produce a "better" forecast -- a fairly lucky position to be in, if you consider that the NWS does a good portion of the heavy lifting for them. We all rely on these weather forecasts to plan our day-to-day activities. For example, before planning a summer grill out over the weekend, we'll check our favorite weather web site to see whether it's going to rain. Of course, we're always left to wonder: Just how accurate are these forecasts? Plotted below is the accuracy of three major weather forecasting services. Note that a perfect forecast means that, e.g., the service forecasted a 20% chance of rain for 40 days of the year, and exactly 8 (20%) of those days actually had rain. There's some pretty startling trends here. For one, The Weather Service is pretty accurate for the most part, and that's because they consistently try to provide the most accurate forecasts possible. They pride themselves on the fact that if you go to Weather.gov and it says there's a 60% chance of rain, there really is a 60% chance of rain that day. With the advantage of having The Weather Service's forecasts and data as a starting point, it's perhaps unsurprising that The Weather Channel manages to be slightly more accurate in their forecasts. The only major inaccuracy they have, which is surprisingly consistent, is in the lower and higher probabilities of raining: Weather.com often forecasts that there's a higher probability of raining than there really is. This phenomenon is commonly known as a wet bias, where weather forecasters will err toward predicting more rain than there really is. After all, we all take notice when forecasters say there won't be rain and it ends up raining (= ruined grill out!); but when they predict rain and it ends up not raining, we'll shrug it off and count ourselves lucky. The worst part of this graph is the performance of local TV meteorologists. These guys consistently over-predict rain so much that it's difficult to place much confidence in their forecasts at all. As Silver notes: TV weathermen they aren't bothering to make accurate forecasts because they figure the public won't believe them anyway. But the public shouldn't believe them, because the forecasts aren't accurate. Even worse, some meteorologists have admitted that they purposely fudge their rain forecasts to improve ratings. What's a better way to keep you tuning in every day than to make you think it's raining all the time, and they're the only ones saving you from soaking your favorite outfit? For me, the big lesson learned from this chapter in Silver's book is that I'll be tuning in to Weather.gov for my weather forecasts from now on. Most notably because, as Silver puts it: The further you get from the government's original data, and the more consumer facing the forecasts, the worse this bias becomes. Forecasts "add value" by subtracting accuracy. --- ## We can only forecast the weather a few days into the future URL: https://www.randalolson.com/2014/06/21/we-can-only-forecast-the-weather-a-few-days-into-the-future/ Published: 2014-06-21 Categories: data visualization Tags: accuracy, meteorologist, nate silver, prediction, rainfall, weather forecast Randy Olson shows why we can only forecast the weather a few days into the future Another fascinating point from Nate Silver's The Signal and the Noise is where he talks about how far into the future we can forecast weather. It's one thing to forecast what tomorrow's weather will be like, but what about next weekend's weather? Or next month's? Silver provided one chart, with data courtesy of Eric Floehr at ForecastWatch.com, that highlights just how hard it is to forecast weather. I've reproduced that chart below. This chart compares three major weather forecasting methods: Persistence: This method assumes that tomorrow's weather will be a lot like today's weather. Tomorrow's temperature will be today's temperature ± a few degrees. Climatology: Since we have decades of historical weather data, we can average what happened in the past on each day to forecast what the weather will be like. This method assumes that the weather on July 4, 2014 in East Lansing, Michigan will be a lot like the weather in East Lansing on July 4 in all the previous years. Commercial Forecasting: Now that the National Weather Service provides so much data about the current weather, we can simulate the weather down to the molecule and create a model of what the weather is going to be like tomorrow. As we'd expect, persistence forecasting performs pretty terribly. If you just take a look at your local weather for the past week, it's rare for temperatures to follow a linear pattern of rising or falling temperatures for more than a day. Even averaging historical data is consistently off the mark by as much as 7°F. The real winners here are the weather models, which can forecast the correct temperature within 4°F up to 3 days out. But even weather models have their limitations: Any forecasts more than a week out are going to be less accurate than climatological forecasts on average, which we've already established makes for a pretty poor baseline. By a week out, small inaccuracies in the weather models build up exponentially, to the point that the model is predicting temperatures far divorced from reality. This observation leads me to wonder why commercial weather forecasting sites like AccuWeather even bother providing forecasts up to two weeks out, considering we'd be better off just looking at the historical averages at that point. So don't bother looking past the 5-day forecasts on your favorite weather site. More likely than not, their forecasts are wrong. --- ## Dolla dolla bill y'all: Relative volume and value of U.S. currency in circulation by bill denomination URL: https://www.randalolson.com/2014/06/20/dolla-dolla-bill-yall/ Published: 2014-06-20 Categories: data visualization Tags: currency, federal reserve, relative value, u.s. dollars Randy Olson visualizes the relative volume and value of U.S. currency in circulation from 1993 to 2013. Last year, Seth Kadish posted some fancy charts showing the relative volume and value of U.S. currency in circulation. Although his charts looked visually appealing, they were fairly heavily criticized because of his use of nested pie charts, which make it notoriously difficult to compare the relative size of areas. For today's data visualization exercise, I decided to try remaking his charts as stacked area charts instead. This was a fairly straightforward exercise: I took the volume and value data from the Federal Reserve's data repository and remade the charts in Excel. Note that I tried to keep everything else as similar as possible, including the color scheme (for better or for worse). Whereas there have always been more $1 bills in circulation (about 35% of all bills), $100 bills have steadily grown from 13% of all bills in 1993 to 26.7% in 2013. The rest of the bills have remained relatively consistent throughout the years: $5 bills around 7%, $10 bills around 6%, $20 bills around 23%, and $50 bills around 5%. The only notable spike occurred in 1999, when an abnormally large number of bills were created (except $2 bills), seemingly focusing on $20 bills. Interestingly, many of the (presumably older) bills were destroyed in 2000 and 2001, marking the only years on record where we actually saw a decrease in the total number of bills. Bills larger than $100 were excluded because they comprise too small a portion of all bills to be effectively visualized. Above, we see where the value of U.S. currency is concentrated, primarily in $100 bills. Since more and more $100 bills are being produced nowadays (relative to other bills), the total value of all U.S. currency is shifting over into $100 bills (58.5% in 1993 vs. 77.2% in 2013). $20 bills saw the largest decrease in relative value, dropping from 21.7% in 1993 to a paltry 12.9% in 2013. Given that this trend toward more $100 bills doesn't seem to be stopping, it leaves us to wonder if we'll be dropping some of the smaller bills (particularly, $2 and $5) soon enough in favor of a more efficient bill system. --- ## Percentage of Bachelor's degrees conferred to women, by major (1970-2012) URL: https://www.randalolson.com/2014/06/14/percentage-of-bachelors-degrees-conferred-to-women-by-major-1970-2012/ Published: 2014-06-14 Categories: data visualization Tags: bachelor's degree, college major, education, gender gap, graduation rates, STEM, university, usa, women in science Randy Olson charts out the percentage of Bachelor's degrees conferred to women (by major) and makes some startling findings. One oft-cited problem with Computer Science is its glaring gender disparity: In a given Computer Science class, men will outnumber women as much as 8 to 2 (20% women). This stands in stark contrast to most other college majors, which have women outnumbering men 3 to 2 on average (60% women). This observation made me wonder: Are other STEM majors suffering the same gender disparity? To get at that question, I checked into the NCES 2013 Digest of Education Statistics and looked at the gender breakdown from 1970-2012 for every major they report on. I charted the data below to offer a bird's eye view of the trends. You can download the cleaned data set here. Edit: For a perspective on the gender gap in female-dominated majors, please look here. Today's trends The woman-dominated majors of today are unsurprising to anyone who has attended a large university in the U.S.: Health Professions (85% women): nursing assistant, veterinary assistant, dental assistant, etc. Public Administration (82%): social work, public policy, etc. Education (79%): pre-K, K-12, higher education, etc. Psychology (77%): cognitive psychology, clinical psychology, etc. Surprisingly to me, most of the STEM majors aren't doing as bad gender disparity-wise as I expected. 40-45% of the degrees in Math, Statistics, and the Physical Sciences were conferred to women in 2012. Even better, a majority of Biology degrees in 2012 (58%) were earned by women. This data tells me that we don't really have a STEM gender gap in the U.S.: we have an ET gender gap! This ET gender gap has severe consequences. Computer Science and Engineering majors have stagnated at less than 10% of all degrees conferred in the U.S. for the past decade, while the demand for employees with programming and engineering skills continue to outpace the supply every year. Compare this to more woman-dominated majors such as Business and Health Professions, which comprise 1/3 of all college degrees in 2012 when combined. Provided that far more women attend college than men, it seems the best way to meet the U.S.'s growing need for skilled programmers and engineers is to focus on recruiting more women -- of any race or ethnicity -- into Computer Science and Engineering majors. The big question, of course, is "How?" With the constant issues of subtle (and sometimes not-so-subtle) discrimination against women in these male-dominated majors, we have quite a tough task on our hands. Looking at the historical trends, maybe we have something to learn from Architecture and the Physical Sciences, given that they were in our position only 40 years ago. Historical trends Perhaps the more fascinating trend in the above graph is how the gender composition of these majors have changed in the past 40 years. Several majors, such as the Health Professions and Education, have been woman-dominated as far back as we have reliable data. But other majors, such as Psychology and Communications/Journalism, didn't see their rise to preference until the late 1970s. Perhaps the most dramatic gender composition change occurred in Agriculture, which started as a gentleman's club in 1970 (only 4% of degrees conferred to women) and grew to an even 50%-50% split by 2012. Going back to Computer Science, we see a rather sad story unfold. The computer scientists of today find themselves in the same disposition as the computer scientists of the 1970s: Only ~15% of the CS degrees were conferred to women. Then the late 1970s and early 1980s finally looked promising: With a peak at 37% of all CS degrees in 1983, it seemed as though Computer Science might join the rest of the majors with a more even gender distribution. But 1984 saw fewer women graduating with a CS degree, and the trend has followed a downward spiral ever since. What was it about the 1970s and early 1980s that made Computer Science more welcoming to women? And what changed? --- ## Football referees are unfair when awarding penalty kicks URL: https://www.randalolson.com/2014/06/13/football-referees-are-unfair-when-awarding-penalty-kicks/ Published: 2014-06-13 Categories: data visualization Tags: biased decisions, decision making, football, penalty kick, referee, soccer Randy Olson explains why football referees are unfair when awarding penalty kicks. After the controversial referee call that decided yesterday's World Cup football match between Brazil and Croatia, I wondered if professional football referees had any consistent biases when they make these potentially controversial calls. After all, referees are forced to make match-altering decisions in just minutes, so it's highly likely that these snap decisions are affected by inherent human biases. It just so happens that one German researcher, Wolf Schwarz, published a report on this topic a few years ago. The results are bound to shock you. Croatian players argue with the referee over his controversial penalty kick call in the 2014 Brazil vs. Croatia World Cup match If referees awarded penalty kicks in a completely unbiased manner, we would expect the distribution of matches with penalty kicks to follow a Poisson distribution: Most matches with no penalty kicks, a good number of matches with 1 penalty kick, a handful of matches with 2 penalty kicks, and so on. Instead, Schwarz found a peculiarity in the data: There are far more 2-penalty kick matches than would be expected by chance. It appears that referees have a bias to turn some of these 1-penalty kick matches into 2-penalty kick matches. The burning question is: Why? To explore this question a little further, Schwarz looked at the breakdown of which teams are awarded the penalty kicks. Surprisingly, the Home team has a considerable advantage when it comes to penalty kicks: 70.6% of all penalty kicks were awarded to the Home team. So not only do Home teams have the advantage of more fans cheering them on, but it appears that the referees are on their side too. Using the above breakdown as a baseline, Schwarz then focused in on the 2-penalty kick matches to see what was going on. Could it be that referees are trying so hard to be fair to both teams that they're trying to give both teams an even number of penalty kicks? The data seems to suggest that this is the case: If the Home team received the first penalty kick, then the Away team received the second penalty kick 48% of the time -- a marked increase from the baseline 29.5%. Similarly, if the Away team received the first penalty kick, then the Home team received the second penalty kick 92.5% of the time -- an incredible display of referee bias. But what if the first penalty kick wasn't converted? Would the referee still feel pressured to award the other team a penalty kick to "even things out"? When Schwarz split the data up by whether the first team converted its penalty kick (here, for the Home team only), it seems the answer is "no." If the first team didn't convert its penalty kick, referees awarded the second penalty kicks as normal (68.4% Home, 31.6% Away). But when the first team did convert its penalty kick, the referee bias popped up again (46.5% Home, 53.5% Away). Clearly, professional football referees are trying so hard to be fair when awarding penalty kicks that they're actually unfair in making their calls. These findings could have sweeping implications for professional football, and especially the World Cup going on this week, if players start taking advantage of this referee bias to score an easy penalty kick. Schwarz looked at several other aspects of the referee bias in his research paper, for example showing that the second penalty kick is awarded much faster than the first penalty kick, which indicates that referees are lowering their standard for what deserves a penalty kick after they awarded the first one. I highly recommend giving his paper a read. --- ## College degrees awarded per capita in the U.S.A. URL: https://www.randalolson.com/2014/06/12/college-degrees-awarded-per-capita-in-the-usa/ Published: 2014-06-12 Categories: data visualization Tags: associate, bachelors, college degree, graduation, masters, phd, university, usa Randy Olson visualizes how many degrees have been awarded in the U.S.A. between 1987 and 2010. As I'm reaching the end of my PhD, I've started thinking more about what I'll be doing afterwards. It's my dream job to teach and do research as a professor, but the prospects aren't promising. There's a ton of competition for academic jobs nowadays, much more so than 20-30 years ago. That made me wonder at lunch yesterday: just how much more competition is there? One way to look at that is to look at the number of degrees awarded. A common mistake that people make when visualizing the number of degrees awarded each year is that they just report the raw number. That's not very useful if you want to compare the number of degrees awarded between years (say, to 20 years ago), because as time goes on, our population grows larger. The larger the population is, the more people there are getting degrees. Therefore, it's important to scale these measures per capita to make a fair comparison. Below, I plotted the number of degrees awarded per capita for four types of degrees: PhDs, Master's, Bachelor's, and Associate degrees. I gathered the degree data from the IPEDS database, and the U.S. population estimates from WorldBank. I'll discuss each plot in detail below. PhDs awarded was the most interesting to me, because these are the folks I'm competing for a job with. 20 PhDs per 100k people means that only 0.02% of the population earns a PhD every year, truly a rare feat that only the most dedicated students can accomplish. Yet here's the stickler: At least in biology, less than 10% of those graduates end up in academic faculty positions, whereas 50% of them share my dream of becoming a university professor. There's incredibly intense competition for an academic faculty job even with such a small set of qualified candidates! The dip in PhDs awarded in the early 2000s stands out here, but I can only guess at the underlying cause. What happened in academia in the mid-late 1990s that made fewer people pursue a PhD? In contrast to the other degree types, Master's degrees have been on a steady incline for the past 20+ years. Interestingly, this means that the recent surge of young students lining up to get a Bachelor's degree haven't been following up to get a Master's -- at least, any more than they would've prior to the college degree surge. What's really crazy here, though, is that after 23 years, we're producing twice as many Master's degrees per capita than we used to. That really raises the question of whether the U.S.A. is really in an education crisis that it has been purported to be. Bears, beets, and Bachelor's degrees! The numbers are really staggering for this plot. About 0.54% of the U.S. population earns a Bachelor's degree each year. And that number is still growing rapidly. At this rate, I wouldn't be surprised if 1% of the U.S. population is earning a Bachelor's degree per year by 2020. An interesting facet of this plot is that there's a dip in Bachelor's degrees awarded in the mid-late 1990s. Perhaps this could explain the dip in PhDs awarded in the early 2000s: Fewer people were graduating from college in the mid-late 1990s, so there were fewer people to follow up and graduate with a PhD 5-7 years later. Associate degrees have followed the same trends as the other degrees, even with the lost interest in them in the 1990s. I'd really like to know what caused the general disinterest in college degrees in the 1990s; could it have been related to the Dot-com bubble? By popular demand, I plotted all of the degree types below (now including Professional degrees) so all degree types can be compared in a single chart. Here we again see the explosive growth in popularity of Master's degrees, surpassing the growth of even Bachelor's degrees. --- ## The key to Magnus Carlsen's success as a chess grandmaster URL: https://www.randalolson.com/2014/06/07/the-key-to-magnus-carlsens-success-as-a-chess-grandmaster/ Published: 2014-06-07 Categories: analysis, data visualization Tags: chess games, grandmaster, magnus carlsen, world chess champion Randy Olson explains the key to Magnus Carlsen's success as a chess grandmaster. For the fifth installment of my series of posts analyzing a data set of over 650,000 chess tournament games ranging back to the 15th century, I wanted to focus in on Magnus Carlsen and try to understand what makes him such an exceptional chess player. In November 2013, Magnus Carlsen soundly defeated the reigning World Chess Champion, Viswanathan Anand, and added "World Chess Champion" to his wardrobe of prestigious titles that he's earned in the past decade of playing chess. Many seem to think that it's Carlsen's ability to wear his opponents down and win games that's been the key to his success lately, but a recent analysis of his games since 2001 seems to suggest otherwise. Anand vs. Carlsen in the 2013 World Chess Championship Magnus Carlsen's meteoric rise to the top started like many of the other chess prodigies. He appeared on the scene in the early 2000s as an ambitious young grandmaster, quickly rising to the ranks of the elite (2750+ Elo rating) within half a decade. Viswanathan Anand, Vladimir Kramnik, and Veselin Topalov all followed similar trajectories in the 1990s, but eventually stagnated in the sub-2800 territory under the shadow of Garry Kasparov. In 2009, Carlsen seemed to be suffering the fate of the many prodigies before him. Then in 2010, Carlsen made an incredible move: After winning a flurry of chess tournaments, he surpassed Kasparov's Elo rating and never stopped growing from there. At this rate, Carlsen seems poised to be the first ever chess player to break a 2900 Elo rating. What has been the key to Carlsen's success? Below, I charted Carlsen's win, loss, and draw rates since 2001. Surprisingly, Carlsen doesn't seem to be winning more games today than he did in 2001. Instead, it appears the key to his success is taking games that he used to consistently lose -- especially games as Black -- and instead forcing them into a draw. Interestingly, this trend mimics the evolution of chess game outcomes since 1850: Every year, more games are ending in draws rather than conclusive wins. It seems that the ideal modern grandmaster is better at forcing a draw to prevent a loss than checkmating their opponent. --- ## Fastest growing and declining surnames in the U.S. URL: https://www.randalolson.com/2014/05/31/fastest-growing-and-declining-surnames-in-the-u-s/ Published: 2014-05-31 Categories: data visualization Tags: ethnicity, race, surname trends, U.S. Census, usa Randy Olson charts the fastest growing and declining surnames in the U.S. according to the U.S. Census. Since my teens, I've been curious how many Olsons there are out there. I've seen some "celebrity" Olsons and even made friends with a couple Olsons throughout my life, but I've always assumed that Olson is a fairly rare last name. To my glee, just yesterday I discovered that the U.S. Census Bureau releases lists of the most common surnames in the U.S. every 10 years. I was quick to look up my surname: 163,502 Olsons in 2000, up +19.5% from 1990. The Olsonian empire is growing! If you want to look yours up, here's the raw data: 1990 | 2000 Of course, I couldn't stop there. Below, I plotted the fastest growing and declining surnames in the U.S. from 1990-2000. If you take note of the racial/ethnic background of the surnames, you'll see an eye-opening trend: Hispanic/Latino surnames are rapidly growing, whereas White/Black surnames are steadily declining in the U.S. (I grouped White/Black because most of the declining surnames were evenly split between White/Black.) For those who follow the news, this finding should be fairly unsurprising: Hispanics/Latinos have been leading the U.S. in population growth for quite some time now. So while Lopez is probably gonna be alright, Jackson will soon have to beat it from the most common surname charts. Just for fun, here's the growth trends for some celebrity surnames: Roberts: 366,215 in 2000 (-3.8% from 1990) Ford: 178,397 (-12.5%) Carey: 54,924 (+16.2) Monroe: 53,475 (-2.3%) Hanks: 17,141 (+14.9%) Pitt: 8,666 (-12.9%) Eastwood: 5,113 (+2.8%) Bieber: 4,294 (-13.7%) Cruise: 3,058 (-38.5%) Johansson: 2,429 (-2.3%) The U.S. Census Bureau hasn't released the list for 2010 yet, but as soon as they do, I'll update this post. --- ## A data-driven exploration of the evolution of chess: Moves, captures, and checkmates URL: https://www.randalolson.com/2014/05/27/a-data-driven-exploration-of-the-evolution-of-chess-moves-captures-and-checkmates/ Published: 2014-05-27 Categories: analysis, data visualization Tags: capture piece, checkmate, chess game, chess move, chess theory, history of chess Randy Olson analyzes over 650,000 chess tournament games to explore how common chess moves have changed over time. For the 4th installment in my series of blog posts exploring a data set of over 650,000 chess tournament games ranging back to the 15th century, I wanted to look at how chess moves have changed over time. Again, I only have reliable data on chess games back to 1850, so 1850 will be my starting point. One thing I was interested in is whether preferences for specific chess moves have changed over time. Was the all-powerful Queen more popular in the past, then lost favor as new strategies developed? Or has capturing pieces become more common nowadays than in previous years? Thankfully, each chess game is recorded in PGN format, which means that it stores every move each player made, the outcome of the game, etc. Here's an example game in PGN format: [Event "F/S Return Match"] [Site "Belgrade, Serbia Yugoslavia|JUG"] [Date "1992.11.04"] [Round "29"] [White "Fischer, Robert J."] [Black "Spassky, Boris V."] [Result "1/2-1/2"] 1. e4 e5 2. Nf3 Nc6 3. Bb5 a6 {This opening is called the Ruy Lopez.} 4. Ba4 Nf6 5. O-O Be7 6. Re1 b5 7. Bb3 d6 8. c3 O-O 9. h3 Nb8 10. d4 Nbd7 11. c4 c6 12. cxb5 axb5 13. Nc3 Bb7 14. Bg5 b4 15. Nb1 h6 16. Bh4 c5 17. dxe5 Nxe4 18. Bxe7 Qxe7 19. exd6 Qf6 20. Nbd2 Nxd6 21. Nc4 Nxc4 22. Bxc4 Nb6 23. Ne5 Rae8 24. Bxf7+ Rxf7 25. Nxf7 Rxe1+ 26. Qxe1 Kxf7 27. Qe3 Qg5 28. Qxg5 hxg5 29. b3 Ke6 30. a3 Kd6 31. axb4 cxb4 32. Ra5 Nd5 33. f3 Bc8 34. Kf2 Bf5 35. Ra7 g6 36. Ra6+ Kc5 37. Ke1 Nf4 38. g3 Nxh3 39. Kd2 Kb5 40. Rd6 Kc5 41. Ra6 Nf2 42. g4 Bd3 43. Re6 1/2-1/2 Notice the extra notation beside the location that the piece moved to: N, B, x, +, etc. Each of these symbols have a particular meaning, e.g., "Qxe7" means that the Queen moved to e7 and took a piece. This notation makes it fairly easy to parse out what pieces are moving where, how many pieces were captured, etc. I charted the evolution of preferences for these movements below. Piece captures One of the biggest questions I wanted to answer with this project is whether capturing pieces has become more or less common nowadays. From my own experience, I noticed that as I became more skilled at chess, I became less focused on capturing all of my opponent's pieces and more focused on controlling the board. Has chess as a sport similarly progressed this way over time? The chart below shows the rate at which pieces were captured over time. "0.2" means that a piece was captured every 1 / 0.2 = 5 ply. Over time, the average chess game has consistently ended with about 16 pieces captured between the two sides. Despite the fact that chess games are getting longer, more pieces aren't being captured in that extended time period. Whereas a piece was captured every 4 ply in 1850, a piece is captured every 5 ply in 2014. This may indeed be because chess games are increasingly becoming more strategic, focusing on gaining control of the board rather than capturing more pieces. Checkmates If chess games are becoming more strategic, then we should expect to see more checkmates over time. Surprisingly, we see the opposite: Less than 2% of expert chess games end in a checkmate in 2014, down from 8% in 1850. I was puzzled by this finding until it occurred to me that most expert chess players are able to predict the next few moves in every game, and therefore resign or call a draw well before the checkmate has occurred. The overall decline in checkmates over time is possibly explained by the fact that draws are becoming the norm in expert chess play, meaning there's fewer games that even come close to a checkmate. Piece preferences We can also look at whether chess players have preferred to use different pieces over time. My favorite piece when I first started to play was always the Queen, but in more recent games I've discovered how powerful Knights can be early on. Below, I charted out the move rates of the pieces over time. "0.33" means that the piece is moved every 1 / 0.33 = 3 ply. Rooks became considerably more popular to use between 1850 and 1900, then leveled off at being used every 7 ply since then. I'd love to hear a chess historian's perspective on this. Perhaps a popular chess book was published around 1850 highlighting new strategies for the Rook? Meanwhile, Knights and Pawns saw a steady increase in popularity from 1920-1970. My best guess is that this is a direct result of the rise of the Hypermodernism school of chess after WWI, which advocated usage of Pawn chains and Knight outposts, both strategies heavily involving Pawns and Knights. Interestingly, Pawns and Knights started falling out of favor in the 1970s, right when Bobby Fischer shook up the chess world with his dramatic march to become World Champion. Did Fischer's breathtaking campaign change chess as we knew it? Finally, King, Queen, and Bishop move rates have remained more-or-less the same over time. I actually find it pretty incredible that we can see much of a change in piece preferences at all, considering that chess strategies are changing so much over time. Castling Castling has been a very popular move throughout chess history. Over 1850-2014, only 16% of the games had one player not castle, and less than 4% of the games had both players not castle. (Surprisingly, many of those games lasted more than 50 ply!) Kingside castling is by far the most popular move (used by 80% of players today) because it requires only 2 pieces to be moved out of the way before it can be done, instead of 3 pieces with the queenside castle. Considering that preferences for queenside castling (below) have remained fairly consistent, I can only guess that players who previously never castled started to realize the value of castling in the 1890s. In contrast to the kingside castle, players have consistently used the queenside castle only about 8% of the time from 1850-2014. It seems that expert chess players want to get their Rooks into play as quickly as possible, leaving the queenside pieces for play later in the game. That's all for today. In the next installment, I'll be looking at preferences for specific locations on the board. --- ## Celebrating 2 years of research blogging by analyzing my blog URL: https://www.randalolson.com/2014/05/27/celebrating-2-years-of-research-blogging-by-analyzing-my-blog/ Published: 2014-05-27 Categories: analysis, data visualization Tags: blogiversary, reddit, research blogging, traffic analysis Randy Olson celebrates his blogiversary by analyzing 2 years of web traffic to the blog. In May 2012, I started this blog to rave about the IPython Notebook, a new scientific computing tool that's still an integral part of my research workflow today. Two years have passed, and I've written about a breadth of topics ranging from statistics tutorials to science outreach to my PhD research to evolution to chess... and even the world's deadliest actors. This blog has proven to be an incredibly useful outlet for fleshing out ideas and projects that I never would've finished without it, and even better, it's connected me with hundreds of brilliant folks who have given me thoughtful feedback about my writing and projects. Coincidentally, this blog's 2nd birthday was also marked by another milestone: It just hit 1,000,000 pageviews! Given this site's recent reputation as a data analysis blog, I figure the only appropriate way to celebrate its birthday is to analyze the blog itself and how it reached this milestone. Without further adieu... Where did all that traffic come from? My blog only has about a dozen subscribers, so it's not like I have a dedicated following checking in every day. 90% of my traffic is referrals from other web sites, the majority from reddit. In the graphs below, I broke my traffic down into a few categories of my largest traffic drivers, with news outlets, other blogs, etc. grouped into the "Other" category. The first graph excludes reddit so you can see the early trends. My blog went mostly ignored for the first 5 months, yet for some reason I kept writing and publishing blog posts to an empty audience. I set up a Search Engine Optimization (SEO) widget in September 2012, which had an immediately noticeable effect on my search engine traffic through Google. Search engines have been a reliable source of traffic over the years, so bloggers to take note: If you want to be found, set up SEO on your web site! Some of my blog posts started getting picked up on various news outlets in late 2013, which led to a bump in the "Other" category. Twitter and Facebook have been a decent source of traffic over the years, but even they've been outpaced by search engine traffic. (And I have a decent-sized Twitter following to tweet to!) And all of them pale in comparison to... ... reddit. I've been an active redditor for over 3 years now, and I've been posting my blog posts there whenever I found an appropriate subreddit to see if anyone finds them interesting. Ever since I started blogging about my data visualization work, the traffic from reddit has exploded, some days reaching 115,000 pageviews in a single day. Social media folks are always talking about how much traffic Twitter, Facebook, etc. can drive, but my experience has always been that reddit is much better for finding people to share and discuss your work with. Another phenomenon you'll notice is that spikes in reddit traffic lead to spikes in the other source's traffic shortly thereafter. The spike in reddit traffic in early January was from a couple of my "deadliest movies" blog posts reaching the front page of /r/movies and /r/dataisbeautiful. The day after, news outlets picked up on the posts and started writing about them, and that's when Twitter and Facebook started to pick up on them. reddit seems to be a springboard for new, original content that wouldn't be found without it -- which is why I love reddit so much. (Thanks for giving me a voice, reddit!) Who reads the blog? I don't have much information about my blog's readers, but here's what I have. Geographic location Half of my blog's visitors are from the USA, leaving the other half to be filled by the rest of the world (mostly Canadians and Europeans). It's been a goal of mine to have at least one visitor from every country, and as you can see, I only have a few countries left to fill. I'm especially proud of the fact that someone from Svalbard (population = 2,642) visited my blog. Internet speed by country Given that my blog has had visitors from all over the world, it's a fun exercise to compare who has the fastest internet speeds. Below are the top 10 fastest and slowest countries that had visitors to my blog. I only included countries that had at least 1,000 visitors to my blog. I'm pretty shocked that the USA doesn't show up in the top 10 fastest, given that my blog is hosted in the USA. Countries with the fastest internet connections Countries with the slowest internet connections Browser popularity and speeds This analysis would not be complete without a comparison of browser popularity and speeds. You can tell that my blog attracts a more tech-savvy crowd, given that over half of its visitors use Chrome. Despite Chrome's claim to being one of the fastest browsers out there, it's been left in the dust by Safari of all browsers! And poor Internet Explorer still can't keep up with the rest of the modern browsers. OS popularity Even though its native browser didn't hold up in the browser race, Windows is still the dominant OS out there. But notice how Apple products are catching up with Windows in terms of market share? What are the most popular posts? It's hard to believe that I've published over 60 blog posts over the past 2 years (an average of 2.5/month!). Below is a list of the most popular ones, sorted by pageviews. 114,031 - A data-driven exploration of the evolution of chess: Popularity of openings over time 106,683 - It's impossible to work your way through college nowadays 99,603 - Top 25 most violence packed films of all time 72,889 - Top 25 deadliest actors of all time by on-screen kills in movies 69,513 - Retracing the evolution of Reddit through post data 47,106- Statistical analysis made easy in Python with SciPy and pandas DataFrames 44,919 - Top 25 most murderous directors of all time 44,911 - Programming Language Breakdown for the HealthCare.gov Website 43,427 - Chess tournament games and Elo ratings 41,592 - It's impossible to work your way through college nowadays, revisited with national data Thanks to everyone for your support over the years. Here's to another 2 years! --- ## A data-driven exploration of the evolution of chess: Popularity of openings over time URL: https://www.randalolson.com/2014/05/26/a-data-driven-exploration-of-the-evolution-of-chess-popularity-of-openings/ Published: 2014-05-26 Categories: analysis, data visualization Tags: chess game, chess theory, history of chess, opening Randy Olson analyzes over 650,000 chess tournament games to explore how chess openings have grown and waned in popularity over time. For the 3rd installment in my series of blog posts exploring a data set of over 650,000 chess tournament games ranging back to the 15th century, I wanted to look at how chess openings have grown and waned in popularity over time. Again, I only have reliable data on chess games back to 1850, so 1850 will be my starting point. The first few moves of a chess game, known as the chess opening, are one of the most-studied aspects of the game, largely because of how important they can be. If you don't start off with a good opening, you could doom yourself to defeat before the game really even begins. It's therefore no surprise that one of the key steps to becoming a skilled chess player is studying and memorizing the many varieties of openings. Hundreds of openings have been developed since 1850, so it should make for an interesting exercise to see how these openings have evolved since then. Each chess game is recorded in PGN format, which means that it stores every move each player made, the outcome of the game, etc. Here's an example game in PGN format: [Event "Hoogovens A Tournament"] [Site "Wijk aan Zee NED"] [Date "1999.01.20"] [EventDate "?"] [Round "4"] [Result "1-0"] [White "Garry Kasparov"] [Black "Veselin Topalov"] [ECO "B06"] [WhiteElo "2812"] [BlackElo "2700"] [PlyCount "87"] 1. e4 d6 2. d4 Nf6 3. Nc3 g6 4. Be3 Bg7 5. Qd2 c6 6. f3 b5 7. Nge2 Nbd7 8. Bh6 Bxh6 9. Qxh6 Bb7 10. a3 e5 11. O-O-O Qe7 12. Kb1 a6 13. Nc1 O-O-O 14. Nb3 exd4 15. Rxd4 c5 16. Rd1 Nb6 17. g3 Kb8 18. Na5 Ba8 19. Bh3 d5 20. Qf4+ Ka7 21. Rhe1 d4 22. Nd5 Nbxd5 23. exd5 Qd6 24. Rxd4 cxd4 25. Re7+ Kb6 26. Qxd4+ Kxa5 27. b4+ Ka4 28. Qc3 Qxd5 29. Ra7 Bb7 30. Rxb7 Qc4 31. Qxf6 Kxa3 32. Qxa6+ Kxb4 33. c3+ Kxc3 34. Qa1+ Kd2 35. Qb2+ Kd1 36. Bf1 Rd2 37. Rd7 Rxd7 38. Bxc4 bxc4 39. Qxh8 Rd3 40. Qa8 c3 41. Qa4+ Ke1 42. f4 f5 43. Kc1 Rd2 44. Qa7 1-0 With a bit of text parsing, I can count the number of times each chess opening was used on a per-game basis. For this analysis, I'll look at the openings in four classes: White's first move, Black's first move, White's second move, and Black's second move. White's first move It's a well-known fact that White has a small advantage at the beginning of the game. To maintain this advantage, White should press their advantage to take over the middle of the board as quickly as possible. The most popular first White moves from 1850-2014 are shown below. Note that all of these are fairly aggressive openings that build toward control of the middle of the board. In 1850, White openings were fairly homogeneous: Most chess experts played King's Pawn. Chess players didn't begin to explore variants of the King's Pawn in earnest until the 1890s, when Queen's Pawn (moving a Pawn to d4) started to replace King's Pawn in some player's repertoires. The 1920s saw another burst of innovation with the rising popularity of the Zukertort Opening (moving the Knight to f3) and the English Opening (moving a Pawn to c4), which completed the set of staple first-turn openings that are really ever used nowadays. Black's first move Many of Black's opening moves are more defensive in nature and attempt to undermine White's initial advantage. In 1850, it was standard fare for Black to match the ever-popular King's Pawn by moving a Pawn to e5 (the Open Game). Although I typically group unpopular openings into the "Other" category, I wanted to point out the short-lived spike in popularity of the Pirc Defence in the 1850s. Though the Pirc Defence is typically thought of as a relatively new opening, Moheschunder Bannerjee used this opening almost exclusively in his 50+ games against John Cochrane, winning 40% of the games (far above his overall 24% win rate as Black). Moreover, the rise of the Queen's Pawn in the 1890s resulted in the rise of the Closed Game in the 1890s. Black openings similarly saw a burst of innovation in the 1920s, with the development of the Indian Defence in response to the Queen's Pawn, and the introduction of the ever-popular Sicilian Defence in response to the standard King's Pawn. By 2014, the Open Game is well past its glory days, and seems to be on its way out. The French Defence seems to have been a staple Black opening for the past 164 years, consistently comprising 5%-10% of all chess games. Amusingly, the French Defence has a reputation for solidity and resilience, which is also reflected in its historical usage. White's second move Here's where things get complicated. I noted in the first section that the most popular first moves for White have historically been King's and Queen's Pawn, so that's why the more popular second moves for White exclusively start with them. The Zukertort and English Openings simply haven't become popular enough yet for their followup moves to show up here. With the waning popularity of the Open Game over time, it's no surprise that the responses to it have similarly declined. By 2014, the typical response to the Open Game is to play the King's Knight, with the once-popular King's Gambit and Vienna Game becoming all but extinct. The Sicilian Defence's explosive rise to popularity is again reflected here, with the Open Sicilian (Knight to f3) becoming White's standard response. Again, White's response to Black's French Defence (moving a Pawn to d4) has remained consistently popular over time, rarely dropping below 5% of the games played each year. To avoid being overly wordy here, I'll allow the visualization to speak for itself and leave the reader to explore the remaining trends as they please. Black's second move If you're familiar with chess, you know how quickly the set of possible moves grows with each move a player makes. After White and Black's first turn, the board will be in one of 400 unique positions. After their second turn, there are 197,742 possible positions. And after only 3 turns, 121 million possible positions. This means that if you play enough chess, it's highly likely that you will play a game that no one has ever played in the history of our universe. You can only imagine how difficult it would be to visualize all possible chess moves even up to the third turn. Despite the infinite possibility in chess, there appears to be a strong bias toward a small subset of openings. In this data set, there were roughly 4,000 unique openings, and the 30 most popular ones comprise 70% of all chess games. Below is a visualization of the distribution of those 30 most popular openings from 1850-2014. (Have any thoughts on a better way to visualize this data? Please leave them in the comments! I've already reached the limit of what area charts can effectively visualize by Black's second move.) Interestingly, chess appears to be becoming more diverse over time. Whereas there were less than 100 unique openings by the end of both player's second turn in 1850, there were over 1,000 unique openings by 2014. This may be an artifact of the data set, however, because there are far more games recorded in the 21st century in this data set. That's it for today. In the next installment, I'll be looking at more higher-level features of player strategy over time. --- ## A data-driven exploration of the evolution of chess: Game lengths and outcomes URL: https://www.randalolson.com/2014/05/24/a-data-driven-exploration-of-the-evolution-of-chess-match-lengths-and-outcomes/ Published: 2014-05-24 Categories: analysis, data visualization Tags: chess game, draw, expert, first-move advantage, history of chess, stalemate Randy Olson analyzes over 650,000 chess tournament games to explore how chess has evolved over time. For the second in my series of blog posts exploring a data set of over 650,000 chess tournament games ranging back to the 15th century, I wanted to look at how chess has changed over time. Nobility and scholars alike have played chess for over 1500 years, and chess has changed considerably since its inception in the 6th century AD. Since I only have reliable data on chess games from 1850-2014, I'll start this analysis at 1850. Chess has been revolutionized several times since 1850. 1851 marked the first international chess tournament in London, leaving the German Adolf Anderssen as the official best chess player in Europe at the time. The 20th century saw several breakthroughs in chess theory as chess players began to treat chess as a science more than a pastime. With the advent of computers in the mid-1900s, chess players started analyzing games and writing computer opponents to hone their craft. Then in the 1990s, the widespread adoption of the Internet allowed players to play chess games with anyone in the world online. Magnus Carlsen represents the newest breed of chess players to revolutionize the chess world. That leaves us to wonder: How has chess changed in that timespan? In this post, I'll look at game lengths and outcomes over time. In future posts, I'll look at how openings and strategies have grown and waned in popularity over time. Distribution of recorded chess games We'll start again with some diagnostics. Unsurprisingly, this data set contains far more games from the past 20 years than for the rest of time. It's becoming easier and easier to keep long-lasting records of chess games now, so we can only expect this trend to continue. Sadly, this means that many games in the 20th century and earlier are lost to us -- but we'll work with what we have. Despite these shortcomings, this data set includes many of the most famous games in chess history, including The Immortal Game and Fischer's Game of the Century. Chess games are getting longer The first thing I wanted to look at is whether games have changed in length. My assumption was that due to their extra practice with computers and solid training in chess theory, modern chess players would be much more efficient at closing a game early. The data shows the exact opposite: 21st century chess games are longer than 19th century games. Chess games have in fact steadily become longer since 1970, increasing from 75 ply (37 moves) per game in 1970 to a whopping 85 ply (42 moves) per game in 2014. Furthermore, if the current trend holds, chess games will only keep getting longer as time goes on. (Note: In all of the following plots, the white line is the mean and the shaded blue area is the 95% confidence interval.) This trend could possibly be telling us that defensive play is becoming more common in chess nowadays. Even the world's current best chess player, Magnus Carlsen, was forced to adopt a more defensive play style (instead of his traditional aggressive style) to compete with the world's elite. The first-move advantage has always existed In my previous post, I discovered that the first-move advantage becomes more pronounced the more skilled the chess players are. When we look at the ratio of White:Black wins in non-drawn games over time, we find that there has always been a first-move advantage: White consistently wins 56% and Black only 44% of the games every year between 1850 and 2014. It's quite interesting that despite 150+ years of revolutions and refinement of chess, the first-move advantage has effectively remained untouched. The only way around it is to make sure that competitors play an even number of games as White and Black. Draws are much more common nowadays Since the early 20th century, chess experts have feared that the over-analysis of chess will lead "draw death," where experts will become so skilled at chess that it will be impossible to decisively win a game any more. The plot below seems to support their fears: Only 1 in 10 games ended in a draw in 1850, whereas 1 in 3 games ended in a draw in 2013. The small dip in draws since 1980 looks promising, but it could very well just be noise. Former World Chess Champion José Raúl Capablanca proposed a more complex variant of chess to help prevent "draw death," but it never really seemed to catch on in the tournaments. We're now only left to see whether the computer-aided analysis of chess will push us ever further into a sea of drawn games. So there we have it. This post has given us a high-level look at how chess has evolved since 1850. The first-move advantage has always been an unfair advantage in chess, and chess games are taking longer to conclude and ending in draws more often than 100 years ago. It will be interesting to check in on the state of chess a decade from now to see how these trends hold up. What else can we learn from this data set? Leave your suggestions and explain why it'd be an interesting analysis in the comments. --- ## Chess tournament games and Elo ratings URL: https://www.randalolson.com/2014/05/24/chess-tournament-matches-and-elo-ratings/ Published: 2014-05-24 Categories: analysis, data visualization Tags: chess match, draw, elo rating, expert, first-move advantage, stalemate Randy Olson analyzes over 675,000 chess games to explore the role of Elo ratings. Chess is by far one of my favorite games. Ever since Seth Kadish shared one of his visualizations of square utilization by chess masters, I've been wanting to follow up to see what else we can visualize about chess. A few months ago, Daniel Freeman from ChessGames.com generously opened his chess data set to me to analyze, which contains a huge collection of 675,000+ chess tournament games ranging all the way back to the 15th century. This will be the first in a series of blog posts exploring this data set. To begin, I was interested in Elo ratings and how they predict the outcome of chess games. Chess match: Bobby Fischer vs. Mikhail Tal (1960) The goal of the Elo rating system is to assign a numeric value that represents a player's chess skill. It's a fairly straightforward yet elegant rating system: All new players start at a relatively low Elo rating. If you beat someone, your Elo rating goes up and their rating goes down the same amount, and vice versa if you lose. The number of points your rating changes by is determined by the difference between you and your opponent's rating. For example, if you have an Elo rating of 1600 and beat a 2200 rated player, your ratings are going to change a lot. But if you beat a 1000 rating player (as a 1600 rating player), your ratings won't change much. Therefore, it's in your best interest to play against others around or above your current Elo rating. After dozens of games, you'll eventually arrive at an Elo rating that's representative of your chess skill. Distributions of Elo ratings Let's start by jumping into the diagnostics. Since this is a data set of chess tournament games, most of the rated players have pretty high Elo ratings. The majority of the games I'll be analyzing were played by experts with a 2000+ Elo rating, many in the 2500 range. To give you a sense of what these ratings mean: Bobby Fischer's peak rating was 2785, and Garry Kasparov's was 2851. So we're analyzing games by some pretty talented chess players. Another important factor to look at is the difference in Elo ratings between the two players in each game. Following the wise advice above, most of the chess tournament games were played between competitors with a fairly close Elo rating. This will be important to keep in mind later as I look at the effect of differences in Elo ratings. Enough diagnostics. Let's get into the meat of the data. Elo ratings tend to predict game outcome My experience with playing against higher-rated players is that the game always seems to end quickly with a checkmate. Sure enough, I'm not the only one to experience this: The higher the difference in Elo rating, the faster the game reaches its final conclusion. The typical game with evenly-matched competitors lasts about 90 plies (45 moves), whereas expert vs. rookie games tend to end in a mere 77 plies (38 moves). Some of the shortest chess games end in a Fool's Mate, where Black checkmates White on her second move. I haven't been able to look if such a game happened in this data set, but it's unlikely because Fool's Mates typically only happen in rookie play. In contrast, the longest recorded chess game was between Ivan Nikolic and Goran Arsovic in 1989, who took a whopping 20 hours to play 269 moves that ended in a draw. Talk about dedication! (Note: In all of the following plots, the white line is the mean and the shaded blue area is the 95% confidence interval.) Furthermore, relative Elo ratings strongly predict who's going to win the game. Evenly-matched competitors have a 50/50 chance of winning, but the more uneven the match is, the more it swings in favor of the competitor with the higher Elo rating. I'd imagine the only reason this trend levels out at ~90% is because this data set contains games where a talented new player hasn't quite reached their proper Elo rating yet. As an extreme case, 9-year-old Awonder Liang defeated GM Larry Kaufman in 2012 in a stunning performance that went straight into the world record books. These findings are likely unsurprising for experienced chess players: Elo ratings tend to accurately predict the outcome of a game. For those who are newer to chess, just know this: If you're pitted against someone with a much higher Elo rating, you can expect a quick and decisive defeat. It matters if you're black or white An often-cited problem with chess is the first-move advantage. Because the White player gets to move first, they get to set the stage for how the game will turn out -- if they know what they're doing. Sure enough, we see exactly that effect in this data set: Newer players with an Elo rating drawn games as White. Meanwhile, the first-move advantage becomes increasingly pronounced the more skilled the players are, where expert players win as much as 62% of their non-drawn games when playing as White. Contrary to Michael Jackson's famous pop song, it matters if you're Black or White in chess. Newcomers, take note: Try to play as White as much as possible and learn to harness the first-move advantage. Draws are more common in expert games Finally, I wanted to take a look at the breakdown of wins, losses, and ties by player skill. I've always heard that rookie chess games tend to be more decisive than expert games, but I'd never seen any data on it. The theory makes sense, of course: rookies tend to leave more openings than their more experienced counterparts, making it easier to close the game with a checkmate. Sure enough, we see a dramatic rise in tie games from rookie (20% draws) to expert (50% draws) games. Your eyes aren't fooling you: Half of all expert chess games end in a draw. We also see the first-move advantage in this chart, where it becomes less and less common for players to win as Black in expert games. Several chess theoreticians even argue that in a perfectly played game, the best outcome for Black is a draw. Unsurprisingly, luring your opponent into a draw has become somewhat of an art as Black. So there you have it. Elo ratings predict the outcome of a chess game in several ways. It'll be interesting to break Elo ratings down by year and see how they've evolved over time. What else can we learn from this data set? Leave your suggestions (and why it'd be an interesting analysis) in the comments. --- ## Programming Language Breakdown for the HealthCare.gov Website URL: https://www.randalolson.com/2014/05/22/programming-language-breakdown-for-the-healthcare-gov-website/ Published: 2014-05-22 Categories: data visualization Tags: health care, healthcare.gov, lines of code, programming language, web site Randy Olson breaks down the code base under the HealthCare.gov website. Late last year, the NY Times released an article quoting a specialist working on the HealthCare.gov web site: According to one specialist, the Web site contains about 500 million lines of software code. By comparison, a large bank’s computer system is typically about one-fifth that size. This astronomically large number became the subject of intense criticism over the following months, especially in the wake of HealthCare.gov's initially failed launch. Particularly, a number of software engineering experts brought into question how realistic it is for any software engineering team to even produce a code base that large. Despite this, the 500 million lines of code statistic has been uncritically cited worldwide. Just today, a data visualization poking fun at this statistic made it to the front page of the subreddit /r/dataisbeautiful. Apparently annoyed by this horrendously false statistic for the last time, one programmer on the HealthCare.gov software development team decided to put the statistic to rest. This programmer performed an automated code count for the HealthCare.gov code base and estimated that there it has only about 3.7 million lines of code for the primary code base. Below is the breakdown of programming languages for that 3.7 million lines of code. The programmer clarified: this doesn't include parts of the system used for administrative tasks. and the total number of lines of code controlling the entire system could be anywhere from 5 - 15 million lines of code. So there you go -- as many of us guessed all along, the 500 million lines of code statistic was utterly bogus. Let's share this information and put that bad statistic to rest. --- ## Skyrocketing student enrollment is partly to blame for rising college costs URL: https://www.randalolson.com/2014/05/20/skyrocketing-student-enrollment-is-partly-to-blame-for-rising-college-costs/ Published: 2014-05-20 Categories: analysis, data visualization Tags: american dream, college, education, historical trends, state funding, tuition costs Randy Olson explains why skyrocketing student enrollment is partly to blame for rising college costs. Earlier this month, The Chronicle of Higher Education released an article claiming that student enrollment is partly to blame for rising college costs. I was a little skeptical at first, largely because I'd seen some damning evidence showing that recent state funding cuts were the real culprit. I figured since I have the IPEDS Delta Cost Project Database on hand, I might as well take a look for myself and see what the numbers say. The results were pretty eye-opening. As many of us know, student enrollment has been steadily rising at universities across the U.S. If you take a look at the graph below, the average 4-year public university has seen a ~25% increase in full-time student enrollment between 2000 and 2010. My alma mater, the University of Central Florida, saw its undergraduate student enrollment spike from 19,781 to 37,609 students in that time period -- a 90% increase in only 10 years! Full-time student enrollment in U.S. 4-year public universities What news outlets typically claim lately is that state funding to these public universities has been steadily declining in the past decade, and that these cuts have been the reason we've seen tuition costs rising so rapidly. Their argument isn't entirely incorrect: The average 4-year public university has had its state funding cut significantly in the past decade, especially since the recession in 2008. Michigan State University saw its state budget cut from about $520 million in 2000 to about $380 million in 2010 -- a tremendous cut by anyone's standards. (Note: dollar amounts were adjusted for inflation to 2014 dollars.) But if you'll notice in the chart below, the average 4-year public university is still receiving more state funding in 2010 than it did in the 1990's when adjusted for inflation. So why can't these universities get by if they're actually receiving more state funding than 20 years ago? State funding for U.S. 4-year public universities The answer, of course, is where we started: Student enrollment. If we divide the average 4-year public university's state funding by its enrollment, we get the chart below. Skyrocketing student enrollment has made it so universities have to teach an ever-increasing number of students with an ever-shrinking budget. What's worse is that we're currently at a stage where public universities are receiving the least amount of state funding per student in the past 2+ decades. Faced with the decision to cut spending and deliver lower quality education to more students, or increase tuition and maintain the quality of higher education, it's no wonder that university administrators chose the latter. State funding per student in U.S. 4-year public universities Now, I'd like to reiterate one of my points from my previous posts: College isn't for everyone. We've been selling too many high school students on this New American Dream: If you work long and hard enough, and if you sacrifice enough, you will eventually graduate college without debt and land your dream job. Instead, students are graduating with an ever-increasing amount of debt and finding it hard to find a job related to their degree. Even worse, by encouraging more and more students to go to college, we're exacerbating the college debt crisis and making college more expensive for everyone. Let's face the facts: You don't need a college degree to have a successful career nowadays. There are plenty of lucrative careers that don't even require a Bachelors degree. We need to stop lying to the Millenials by telling them that they need a college degree to get ahead. --- ## Popular subreddits have predictable cycles of virality URL: https://www.randalolson.com/2014/05/16/popular-subreddits-have-predictable-cycles-of-virality/ Published: 2014-05-16 Categories: analysis, data visualization, reddit Tags: front page, hotness, post ranking, reddit, subreddit, virality Randy Olson maps out the cycles of virality on reddit's front page. Ever since I started studying reddit, I've always wondered why the most successful posts were consistently posted around 7-8 am EST. I've had several theories, but until today I wasn't able to back any of them up with real data. After gaining a better understanding of how reddit ranks posts on the front page a couple months ago, I think I've finally figured it out. To get to the default front page of reddit (that most visitors to reddit see), our post needs to get to the top of one of the default subreddits. And to stand a chance of getting to the top of one of the default subreddits, our post needs to make it into the top 25 posts of one of the default subreddits. Otherwise, too few people will see and upvote it enough to push it to the top. Of course, our post is constantly battling with other posts to have a higher hotness score so it can be ranked higher. Therefore, it behooves us to post our link when the other posts in the subreddit have the lowest hotness. In other words, we should strike when our opponents are at their weakest. For the past couple months, I've been measuring the hotness of the front page of the most popular subreddits. This front page hotness score gives us a sense of how easy it will be for our new post to invade a particular subreddit's front page. The higher the score is above 1.0, the more difficult it will be. If the score is below 1.0, then our new post will have a much easier time of getting into the subreddit's top 25. I've plotted the hotness scores by day and hour for several of the popular subreddits below. Notice how there are predictable cycles of virality in several of the popular subreddits: In the mornings to noon EST, the front pages die down and are easier to invade. Around 2 pm EST, the front pages start to pick up again and culminate in the highest virality in the evening -- just in time for the US workforce to get home and browse reddit. It's no wonder, then, that the most successful posts were posted around 7-8 am EST. Around that time, the popular posts from the previous day are cooling down and falling off the front page, leaving an opening for a new post on the block. If these 7-8 am posts can build up enough hotness to reach the subreddit's top 25 by the afternoon, then they're in prime position to get upvoted to the default front page. Not all subreddits follow this trend (especially the more niche ones like /r/dataisbeautiful), but many of the more popular subreddits clearly show that they're dominated by users in the North American time zones. In future posts, I'll follow up with the hotness trends in more niche subreddits. "Breathing" map of virality Just for fun, I made a "breathing" map of virality for 40 of the popular subreddits. Here, each square represents a subreddit and the color of the square represents how hot the subreddit was at the given time. To make the hotness values comparable between subreddits, I normalized each subreddit's hotness scores by the average hotness score for the subreddit. Also, to make this video a little more grandiose, each of the 40 subreddits I'm visualizing are actually copied 40 times into random locations on the grid, making for a total of 1,600 tiles. Red = hotter; blue = colder Methodology I ran a bot between March 23, 2014 through May 11, 2014 that checked the hotness of the front page of several of reddit's most popular subreddits every hour. To determine the hotness of a subreddit's front page, I first computed and summed the hotness of each post in the subreddit's current 25 hottest posts. I then divided the subreddit's summed hotness by the summed hotness of 25 new posts with no upvotes, which is the score I plotted above. The summed hotness of 25 new posts with no upvotes provides a "baseline hotness" where a new post with no additional upvotes would not be able to immediately reach the subreddit's front page. Therefore, new posts in subreddits with a hotness score >= 1.0 will not start on the front page of the subreddit. In contrast, new posts in subreddits with a hotness score < 1.0 are much more likely to make it onto the subreddit's front page. Note the score's small scale; small increases or decreases of the score have a large effect! --- ## Rising and falling subreddits according to virality trends URL: https://www.randalolson.com/2014/05/16/rising-and-falling-subreddits-according-to-virality-trends/ Published: 2014-05-16 Categories: analysis, data visualization, reddit Tags: falling, front page, hotness, post ranking, reddit, rising, subreddit, virality Randy Olson lists several subreddits that are rising and falling in popularity according to virality trends. For the past week, I've been mapping out virality trends on various subreddits to better understand how posts go viral on reddit. As a neat side effect, I discovered a method to detect rising and falling subreddits. Basically, a "rising" subreddit is one that has a large, steady increase in hotness over time (blue to red), whereas a "falling" subreddit is going in the opposite direction (red to blue). Rising subreddits Here's the rising subreddits I've manage to detect: /r/4chan, /r/AlienBlue, /r/AskWomen, /r/battlestations, /r/Demotivational, /r/Diablo, /r/dubstep, /r/Fallout, /r/IWantToLearn, /r/JusticePorn, /r/learnprogramming, /r/MorbidReality, /r/Steam, /r/TrueReddit, /r/woahdude Falling subreddits Here's the falling subreddits I've managed to detect: /r/itookapicture, /r/MURICA, /r/Frugal, /r/Health, /r/Homebrewing, /r/keto, /r/loseit, /r/progresspics, /r/sex, /r/travel --- ## Virality trends in niche subreddits URL: https://www.randalolson.com/2014/05/16/virality-trends-in-niche-subreddits/ Published: 2014-05-16 Categories: analysis, data visualization, reddit Tags: front page, hotness, niche, post ranking, reddit, subreddit, virality Randy Olson maps out virality trends in niche TV show and video game subreddits on reddit. For the past week, I've been mapping out virality trends on various subreddits to better understand how posts go viral on reddit. As a neat side effect, I discovered that several niche subreddits have their own interesting virality patterns. Below are a few that I thought were worth sharing. /r/asoiaf and /r/gameofthrones Every Sunday, a new episode of Game of Thrones airs on HBO. Without missing a beat, the two subreddits dedicated to Game of Thrones light up the day after the show as folks talk about the latest goings-on in the show. This happens pretty consistently for nearly every TV show on reddit. Note that the /r/gameofthrones plot is missing about a week of data. /r/HIMYM and /r/thewalkingdead Just as TV show subreddits rise in popularity as a new TV show comes out, so too do they die out when the series is on hiatus. Since How I Met Your Mother and The Walking Dead ended their seasons in March, /r/HIMYM and /r/thewalkingdead have slowly faded into obscurity... until next season (for /r/thewalkingdead, anyway). Talking about video games isn't so popular on the weekends In the niche video gaming news subreddits /r/Games & /r/gamernews, it looks like gamers sign off of reddit and go back to playing video games on the weekends. --- ## Virality trends in reddit's default subreddits URL: https://www.randalolson.com/2014/05/16/virality-trends-in-reddits-default-subreddits/ Published: 2014-05-16 Categories: analysis, data visualization, reddit Tags: default, front page, hotness, post ranking, reddit, subreddit, virality Randy Olson outlines how the new reddit default configuration has affected the subreddits involved. For the past week, I've been mapping out virality trends on various subreddits to better understand how posts go viral on reddit. It just so happens that the time period I sampled over covered the week that the reddit admins made a huge change to the default set of subreddits. By mapping out the hotness before and after the default set was changed (05/07/2014), we can get a unique look at how the new default configuration has affected the subreddits involved. Previous defaults In a surprising move, the admins removed both /r/AdviceAnimals and /r/bestof from the default list. I won't speculate as to why that happened, but these subreddit's demotion has already had noticeable effects in the past week. /r/AdviceAnimals seems to be slowly going the way of /r/fffffffuuuuuuuuuuuu (i.e., steadily losing virality), and /r/bestof hasn't had a viral day since the change over. New defaults I was lucky enough to already be tracking a handful of the new defaults before they were promoted. Most of the new default subreddits, such as /r/Art, /r/Documentaries, /r/mildlyinteresting, and /r/tifu, saw a huge spike in virality the day they were promoted. /r/dataisbeautiful also saw a initial spike in virality, but seems to be dropping back down to previous levels. I've been moderating /r/dataisbeautiful and we've been keeping a tight reign on what gets posted there, which may explain why we haven't seen a steady rise in virality. The nice part is that we don't seem to have a gap in viral posts on the weekends any more! Whereas /r/DIY, true to its name, seems to have built up most of its virality already before it became a default. Veteran defaults Oddly enough, some of the veteran defaults that remained defaults seem to have been affected by this new setup. Both /r/IAmA & /r/videos have dropped in virality in the past week. It looks like they were already dropping in virality for a couple weeks before the changeover, though, so perhaps their fall is already a foregone conclusion. Meanwhile, most of the veteran defaults, such as /r/funny & /r/pics, have remained relatively unscathed. --- ## U.S. Racial Diversity by County URL: https://www.randalolson.com/2014/04/29/u-s-racial-diversity-by-county/ Published: 2014-04-29 Categories: data visualization Tags: county, ethnicity, map, racial diversity, united states Randy Olson maps out racial diversity in the U.S. The U.S. is typically viewed as a melting pot of races and cultures, but recent maps showing the ethnic distribution of the U.S. seem to hint that the U.S. isn't as well-mixed as we all thought. In this visualization, I mapped out the racial/ethnic diversity of the U.S. to give us a better sense of the hotspots of diversity. To calculate racial/ethnic diversity, I computed the entropy on the "% ethnicity" data for each county using the 6 ethnic categories the U.S. Census tracks: White (non-Latino), African American, Native American, Asian American, Latino, and Other. A county will come out with high entropy when all 6 ethnic categories are as even as possible (i.e., each ~16.7%), whereas it will come out with low entropy if the county is only inhabited by people of one ethnic category. I've included the raw census data if you want to tinker with it yourself. One of the most notable features is that the Midwest and Northeast are fairly homogeneously white. Vermont, New Hampshire, and Maine stand as the pinnacle of racial homogeneity, each with only one or two counties with even a blip of diversity. The only exceptions to this trend are the major cities throughout the U.S., which seem to attract people of all ethnicities regardless of the state the city is in. As a Michigander, I'm the most surprised to see how diverse the Upper Peninsula is. I thought only crazy white people lived up that far in Michigan. Here's the least diverse counties: Tucker County, West Virginia (100% White, non-Latino) Robertson County, Kentucky (100% White, non-Latino) Hooker County, Nebraska (100% White, non-Latino) Hand County, South Dakota (99% White, non-Latino and 1% Latino) Owsley County, Kentucky (98% White, non-Latino and 2% Latino) And the most diverse counties: Aleutians West Census Area, Alaska (31.4% White (non Latino), 5.7% African American, 15.1% Native American, 28.3% Asian American, 13.1% Latino, and 6.4% Other) Aleutians East Borough, Alaska (13.5% White (non Latino), 6.7% African American, 27.7% Native American, 35.4% Asian American, 12.3% Latino, and 4.4% Other) Queens County, New York (27.6% White (non Latino), 17.7% African American, 0.3% Native American, 22.8% Asian American, 27.5% Latino, and 4% Other) Alameda County, California (34.1% White (non Latino), 12.2% African American, 0.3% Native American, 25.9% Asian American, 22.5% Latino, and 5.1% Other) Solano County, California (40.8% White (non Latino), 14.2% African American, 0.5% Native American, 14.3% Asian American, 24% Latino, and 6.2% Other) --- ## If every U.S. state had a surname for bastards, like Game of Thrones, what would each state's name be? URL: https://www.randalolson.com/2014/04/27/if-every-u-s-state-had-a-surname-for-bastards-like-game-of-thrones-what-would-each-states-name-be/ Published: 2014-04-27 Categories: data visualization Tags: asoif, bastard, game of thrones, map, surname, usa Randy Olson maps out the Game of Thrones-like bastard surnames of the USA. There was a fun thread on reddit yesterday that asked users to come up with Game of Thrones-like surnames for every U.S. State. I went through and picked out my favorites from the thread and mapped them below. When reading these, read it as, "You know nothing, Jon ______." --- ## Visualizing evolution in action: Density-dependence and sympatric speciation URL: https://www.randalolson.com/2014/04/17/visualizing-evolution-in-action-density-dependence-and-sympatric-speciation/ Published: 2014-04-17 Categories: data visualization, research Tags: density dependent, evolution, evolutionary processes, fitness landscape, species, sympatric speciation, visualization Randy Olson and Bjørn Østman demonstrate how sympatric speciation can occur via density dependence using fitness landscapes. Fitness landscapes were invented by Sewall Wright in 1932. They map fitness, or reproductive success, of individual organisms as a function of genotype or phenotype. Organisms with higher fitness have a higher chance of reproducing, and populations therefore tend to evolve towards higher ground in the fitness landscape. Even though only two traits can be visualized this way, we can actually observe evolution in action. Building on the idea of fitness landscapes, Bjørn Østman and I decided to create some animations of simulated evolving populations to illustrate concepts of evolution that are typically difficult to comprehend. Here we demonstrate how sympatric speciation can occur when fitness depends on the density of organisms, i.e., density-dependence. Warning: The GIFs on this page are large and may take some time to load. Here's the full video that Bjørn and I submitted to the ALife 2014 Science Visualization Competition. --- ## Visualizing evolution in action: Dynamic fitness landscapes URL: https://www.randalolson.com/2014/04/17/visualizing-evolution-in-action-dynamic-fitness-landscapes/ Published: 2014-04-17 Categories: data visualization, research Tags: dynamic, evolution, evolutionary processes, fitness landscape, visualization Randy Olson and Bjørn Østman use fitness landscapes to demonstrate the effect of dynamically changing environments on evolving populations. Fitness landscapes were invented by Sewall Wright in 1932. They map fitness, or reproductive success, of individual organisms as a function of genotype or phenotype. Organisms with higher fitness have a higher chance of reproducing, and populations therefore tend to evolve towards higher ground in the fitness landscape. Even though only two traits can be visualized this way, we can actually observe evolution in action. Building on the idea of fitness landscapes, Bjørn Østman and I decided to create some animations of simulated evolving populations to illustrate concepts of evolution that are typically difficult to comprehend. Here we demonstrate the effect of a dynamically changing environment on an evolving population. If the environment changes slowly enough, the population can adapt and "keep up" with environmental change. But if the environment changes too quickly, evolution breaks down and the population can no longer adapt to the environment. Warning: The GIFs on this page are large and may take some time to load. Here's the full video that Bjørn and I submitted to the ALife 2014 Science Visualization Competition. --- ## Visualizing evolution in action: Survival of the flattest URL: https://www.randalolson.com/2014/04/17/visualizing-evolution-in-action-survival-of-the-flattest/ Published: 2014-04-17 Categories: data visualization, research Tags: evolution, evolutionary processes, fitness landscape, survival of the flattest, visualization Randy Olson and Bjørn Østman demonstrate the concept of "survival of the flattest" using fitness landscapes. Fitness landscapes were invented by Sewall Wright in 1932. They map fitness, or reproductive success, of individual organisms as a function of genotype or phenotype. Organisms with higher fitness have a higher chance of reproducing, and populations therefore tend to evolve towards higher ground in the fitness landscape. Even though only two traits can be visualized this way, we can actually observe evolution in action. Building on the idea of fitness landscapes, Bjørn Østman and I decided to create some animations of simulated evolving populations to illustrate concepts of evolution that are typically difficult to comprehend. Here we demonstrate the survival of the flattest, a theory stating that when organisms experience a high enough mutation rate, the population will evolve to "flatter" fitness peaks instead of higher peaks. This theory of course flies in the face of the more traditional "survival of the fittest," which would have us think that organisms will always adapt to the highest fitness peak. We hope to show here that there's much more to evolution than "survival of the fittest." Warning: The GIF on this page is large and may take some time to load. Here's the full video that Bjørn and I submitted to the ALife 2014 Science Visualization Competition. --- ## It's impossible to work your way through college nowadays, revisited with national data URL: https://www.randalolson.com/2014/03/29/its-impossible-to-work-your-way-through-college-nowadays-revisited-with-national-data/ Published: 2014-03-29 Categories: analysis, data visualization Tags: american dream, college, education, historical trends, tuition costs, work study Randy Olson explains why it's nearly impossible to work your way through college nowadays. Last weekend, I wrote a brief rant about how it's far more difficult to work your way through college nowadays than 30 years ago. Some folks took it for a scientific study rather than the rant it was, and criticized it for only looking at Michigan State University's tuition trends. In response, I decided to run a proper analysis of national public university tuition data. With the help of some of my awesome Twitter followers, I managed to find a comprehensive data set of the in-state tuition costs for all public 4-year universities in the U.S. from 1987 through 2010. Combining that data with the Federal minimum wage trends from before, we get the chart below showing the number of hours a student would have to work on minimum wage to pay for 1 year of public university tuition in the U.S. To save you the data wrangling, I'll provide the data set here. Hours worked on minimum wage to pay for 1 year of public university tuition in the U.S. We immediately see a trend similar to before, but the data is limited between 1987 and 2010. What about the 1979 and 2013 students, as we previously looked at? To get a better sense of the trend, I fit a linear regression to the data. According to the model, students have to work 23.7 extra hours every year to pay for tuition. If we extrapolate this trend back to 1979 and forward to 2013, we recover the same trend that I found in my previous post: The average university student in 1979 only had to work 182 hours per year (a part-time summer job) to pay for tuition, whereas the average 2013 student had to work 991 hours (a full-time job for half the year). That's over 5x as many hours worked for the same education! I should point out that I'm only considering tuition & fees in this analysis, and I've completely left out room & board, book costs, gas & car repair, and other miscellaneous expenses that build up when students are in college. Given the widespread reports that wages aren't keeping pace with inflation, this plot would only look even more dismal if I factored those costs in as well. Other commenters were eager to point out that I left out financial aid from this analysis. If the Federal aid trends in the past 30 years are any indication, students actually have less of their tuition costs paid for by financial aid nowadays than 30 years ago! With rising costs and lowered financial support, it's no wonder that student debt has spiraled out of control in the past decade. The system is practically setting the modern university student up for financial failure. In summary, I'd like folks to stop toting their college financial success stories as an excuse for the insane costs of tuition nowadays -- you're the exception, not the rule. --- ## It's impossible to work your way through college nowadays URL: https://www.randalolson.com/2014/03/22/its-impossible-to-work-your-way-through-college-nowadays/ Published: 2014-03-22 Categories: analysis, data visualization Tags: american dream, college, education, historical trends, tuition costs, work study Randy Olson explains why it's impossible to work your way through college nowadays. Update (3/29/14): I've written up an analysis of national tuition cost trends in a new blog post. It turns out that Michigan State University's tuition situation isn't uncommon! Earlier today, I ran across a conversation about how the cost of tuition at Michigan State University (MSU) has changed over the years. I had just finished talking with my grandpa over the phone, and he had spent the latter half of the talk extolling the virtues of working your way through college (without family support), so I was rightly annoyed on the topic already. The creator of the discussion pointed to the historical trends for MSU's tuition, and in another comment pointed to the Federal minimum wage trends. If you crunch some of the numbers there, you'll get the chart below showing the number of hours a student must work on minimum wage to pay for a single credit hour at MSU. I'll provide the data set here to save you the number crunching*. Hours of minimum wage work required to pay for 1 credit hour What we see is a startling trend: Modern students have to work as much as 6x longer to pay for college than 30 years ago. Given the reports that a growing number of college students are working minimum wage jobs, this spells serious trouble for any student who hopes to work their way through college without any additional support. Let's crunch a few more numbers to see what a typical year would look like for a student in 1979 and 2013 working her way through college. Most students take 12 credit hours per semester and only attend Fall and Spring semester. That's 24 credit hours per year. The 1979 student would have to work about 10 weeks at a part-time job (~203 hours) -- basically, they could pay for tuition just by working part-time over the Summer. In contrast, the 2013 student would have to work for 35 ½ weeks (~1420 hours) -- over half the year -- at a full-time job to pay for the same number of credit hours. If you've ever attended college full-time, you know that this is basically impossible. Perhaps it's no surprise that tuition costs are rising, and college is becoming less and less affordable by the year. Yet somehow, the idea that we can work our way through college still persists. This ethos seems to be the latest generation's version of American Dream: If you work long and hard enough, and if you sacrifice enough, you will eventually graduate college without debt and land your dream job. But with the way this trend is going, it looks like even long and hard hours at work won't even pay off any more. In short, I'd like my readers to walk away knowing that it's not nearly as easy to work your way through college as it used to be -- stop telling us to do it just because you did a decade or more ago. * Note about the data: For the minimum wage data, I set the minimum wage to the maximum for the year, even if minimum wage was raised in the later parts of the year (e.g. September). For the tuition data, I averaged the Fall and Spring tuition rates (if available), and only used the tuition rates for students admitted that year. I dropped any entries for Summer tuition on the assumption that most students do not attend Summer semester. --- ## The window of virality on Reddit URL: https://www.randalolson.com/2014/03/21/the-window-of-virality-on-reddit/ Published: 2014-03-21 Categories: analysis, data visualization, reddit Tags: front page, hot algorithm, post, reddit, viral Randy Olson visualizes when posts go viral and when they fade into obscurity on Reddit. Have you ever wondered why most of the posts on Reddit's front page are less than 12 hours old? Or why a post with a score of 4,000 is ranked below 3,000 score post? It all has to do with Reddit's window of virality. Each post on Reddit has a score attached to it: score = upvotes - downvotes. Reddit's "hotness" algorithm uses this score in combination with the post's age to rank every single post on Reddit. Amir Salihefendic wrote a fantastic post explaining the nitty gritty of how Reddit's hotness algorithm works, so I won't bother repeating that here. Instead, I'll jump right into the visualization showing us Reddit's window of virality. The y-axis indicates the post's current score; the higher up, the higher the post's score. The x-axis indicates the post's current age; the more to the right, the older the post is. The color indicates the post's "hotness" or virality, with darker shades of red for viral posts and darker shades of blue for posts that don't stand a chance of getting to Reddit's front page. I labeled the dark blue region as "dead" because posts in these regions have the same hotness as a newly submitted post with no upvotes. If our post can be outranked by a post that hasn't even been voted on yet, it doesn't stand a chance of making the front page. Immediately, we see why posts older than 12 hours are such a rarity on Reddit's front page: A 12-hour-old post needs roughly 3x the score to match the hotness of a 6-hour-old post! 18- and 24-hour-old posts don't even stand a chance on the front page unless it's President Obama holding an AMA. Clearly, the life of a viral post on Reddit is short -- so if you make it to the front page, enjoy your precious few hours in the spotlight. It's no surprise then that time plays such a huge role in the success of a Reddit post. If the first 6 hours of a post's life are the most crucial for going viral on Reddit, none of that time can be wasted sitting around when no one's online. In the end, all posts must die. As we see in the top right of this graph, even the most high-scoring posts will be ouranked by a new post with no upvotes at all by the end of their second day of life. Reddit's front page is constantly evolving, and 2-day-old news is just so yesterday. There's of course an important addendum here: This classification really only applies to Reddit's front page, but not individual subreddit pages. If we take a stroll down any of the smaller, non-default subreddits, we'll see plenty of older, sub-500 score posts that are still on the front page of the subreddit. That just means it's easier to get on the front page of the individual subreddits -- but that's a topic for a future post! --- ## The cost of war in Iraq and Afghanistan URL: https://www.randalolson.com/2014/03/11/the-cost-of-war-in-iraq-and-afghanistan/ Published: 2014-03-11 Categories: data visualization Tags: bush, fatalities, obama, operation enduring freedom, operation iraqi freedom, soldier, united states, war on terror Randy Olson visualizes US soldier fatalities from the wars in Iraq and Afghanistan. For the past decade, the folks at iCasualties have been painstakingly keeping records of the fatalities in the wars in Iraq and Afghanistan. Unfortunately, a good portion of the data is presented as tables of numbers that most of us will glaze over by the third row. As Richard Hamming famously said, "The purpose of computing is insight, not numbers." In that spirit, below is my attempt to glean some insights from iCasualties' data through visualization. [caption id="attachment_2793" align="aligncenter" width="584"] Monthly US soldier fatalities in Operation Enduring Freedom. "Operation Enduring Freedom" was the name given to the war in Afghanistan.[/caption] The first 8 years of the war in Afghanistan saw relatively few US fatalities compared to when President Obama sent in several surges of troops between 2009 and 2012. We even see the deadliest month of the war -- August 2011 -- when the Chinook helicopter carrying a team of Navy Seals was shot down in Afghanistan. This plot makes me wonder: if we were able to make progress in Afghanistan without so many fatalities for the first 8 years, were the troop surges really necessary? Also note that there are still US fatalities in 2014. Even though we've moved our gaze to other matters, the war in Afghanistan is still ongoing. [caption id="attachment_2794" align="aligncenter" width="582"] Monthly US soldier fatalities in Operation Iraqi Freedom. "Operation Iraqi Freedom" was the name given to the war in Iraq.[/caption] It's amusing (but sad) to see how far off President Bush's "Mission Accomplished" speech in May 2003 was. Despite his claim that all major combat operations were over in Iraq, the death toll that accumulated over the next 9 years tells us otherwise. One of the first things that stands out to me here are the two deadliest months in 2004. Both of these months align with the deadliest battles of the Iraq war over Fallujah. Note that this was also during election year, when American patriotism and public support for the Iraq war was relatively high. Could these deadly battles have been pushed by politics rather than military strategy? We clearly see the effect of President Bush's troop surges in 2007. Contrary to the troop surges in Afghanistan, after the initial spike in fatalities for the first half of 2007, the fatalities drop and continue dropping for the remainder of the war. Could it be that this troop surge "paid off," considering how deadly the Iraq war was before the surges? If we compare this plot with the plot for Afghanistan, note that 2008 -- conveniently, election year when support for the war was relatively low -- was the only year that didn't see a considerable spike in US fatalities in both theaters. Coincidence? Now it's your turn. What other insights can we glean from these plots? --- ## Harnessing Hashtagify.me and social network analysis to maximize your brand's reach on Twitter URL: https://www.randalolson.com/2014/03/05/harnessing-hashtagify-me-and-social-network-analysis-to-maximize-your-brands-reach-on-twitter/ Published: 2014-03-05 Categories: analysis, tutorial Tags: brand marketing, business intelligence, hashtags, hastagify, social network analysis, twitter Randy Olson explains how to use Hashtagify.me and social network analysis for Twitter. According to an August 2013 PEW report, 18% of all internet users use Twitter on a regular basis. That equates to roughly 500 million people signing into Twitter to check the latest tweets, news, and celebrity gossip every day. It's no surprise that brand marketers have taken an interest in connecting to even a small fraction of those 500 million users with the hope of increasing sales and their brand's reputation. However, like most online social networks, Twitter has proven to be an amorphous entity that even the best social network analysts struggle to understand. How can we make sense of the massive amount of information on Twitter? More importantly, how can we learn from this information to better market our brands on Twitter? In this guide, we're going to walk through how we can use Hashtagify.me and some basic social network analysis techniques to target a specific community, identify the key hashtags for that community, and then properly use the hashtags in our tweets to maximize our brand's reach on Twitter. What can Twitter hashtags do for your brand? Some brand marketers may be wondering why hashtags are even worthwhile. Hashtags take up limited tweet characters and make the tweet awkward to read. Why not focus on writing a clear yet concise tweet that explains the product and why it's worth considering? First and foremost, hashtags are the primary method that users connect to each other on Twitter. If we don't use hashtags, the only users that see our tweets are the ones that follow us on Twitter. Any brand marketer who's used Twitter before can testify how difficult it is to build a loyal following on Twitter. The data about tweet engagement is even more eye opening. Here's a rundown of the relevant points: The more often we tweet, the less likely people are to pay attention to our tweets Shorter tweets (Tweets with hashtags receive 2X more attention than tweets without hashtags Tweets with only 1-2 hashtags receive more attention than tweets with 3+ hashtags This means that if we want to maximize engagement with our tweets, we need to tweet infrequently, be short yet to-the-point, and include 1-2 relevant and popular hashtags in our tweet. That's an awful lot to consider for a 100-character message, isn't it? Fortunately, Hashtagify.me is in the business of making that whole process a breeze. How to target a community using Hashtagify.me As a case study, we're going to focus on the Big Ten sports network. The Big Ten is a collection of 12 large U.S. university's sports programs with hundreds of thousands of fans and followers, and thus represents a prime target for marketing any products or services related to sports. As we work through this case study, imagine how these tools can be applied to market your own brand. To start, we have to identify all of the major hashtags related to the Big Ten network on Twitter. We can start by searching "BigTen" and keeping track of all the other hashtags people use to refer to the Big Ten. After that, we can do the same search process for the names of the universities in the Big Ten and their sports teams. Conveniently, Hashtagify.me provides us with a network visualization of the most related hashtags, so it takes less than 15 minutes to click around through all the hashtags to catch all of the variants on the team names. #Spartans Hashtagify.me Network Shown above, Hashtagify.me also provides a measure of each hashtag's popularity, in other words, how frequently the hashtag is used in relation to all other hashtags. We can also record that popularity measure for each Big Ten-related hashtag and rank the hashtags according to popularity. Popularity of Big Ten sports Twitter hashtags The above analysis is already tremendously helpful for us: Now we know that #Michigan, #Huskers, and #Buckeyes are the three most popular hashtags within the Big Ten community. This tells us that these hashtags probably have the most people paying attention to them because people that use hashtags also generally follow them. Another great feature about Hashtagify.me's hashtag network visualization is that is tells us how related each hashtag is to each other. In this case, "relatedness" is a measure of how often the hashtags are used together in the same tweet. If we again record all of these relatedness scores, we can produce a nice map of the Big Ten hashtag network. There are several different tools that we can use to visualize these kinds of networks, but I used Gephi in this case study because it's free, open source, and easy to use. See Gephi's tutorial on their data laboratory to find out how to input network data and visualize it in Gephi. Big Ten Twitter Hashtag Network Click on the image to go to an interactive version of the network With this network visualization, we now have a better sense of the Big Ten hashtag network. Despite the fact that #B1G is fairly average in popularity for this network, it appears to be highly related to several of the other sports team specific hashtags. Fun fact: What's especially interesting (and puzzling!) here is that despite the fact that Northwestern University, Indiana University, Purdue University, and the University of Iowa have been a part of the Big Ten for over 100 years, their sports teams aren't mentioned very often with other Big Ten sports teams on Twitter. This is in direct contrast with the relatively newer Big Ten teams such as Ohio State University, Michigan State University, and Pennsylvania State University, whose sports teams are all mentioned frequently with other Big Ten sports teams. The causes and implications of this are worth exploring in another post, but let's get back to brand marketing. How to identify key hashtags in a community The most popular hashtags in a community aren't necessarily the best hashtags to use for that community. If we used #Michigan, for example, we'd be reaching a large number of people, but we'd only be targeting people that follow the Michigan sports teams. Instead, we want to identify the hashtag that's both popular and is used frequently with several other popular hashtags in the community. In other words, we want to find the hashtags that are most central to the community. Fortunately, several smart people have already worked out this problem for us. We can use measures such as eigenvector centrality and betweenness centrality to find the most central hashtags. In short: High eigenvector centrality hashtags are the leaders and influencers of the network. High betweenness centrality hashtags are responsible for spreading information between communities in the network. In this case study, we want to find the hashtag that has a large influence on the entire Big Ten community and tends to spread information to all of the other school's sports teams. Just our luck, Gephi has built-in tools to calculate both of these measures, so all we have to do is click a couple buttons and it calculates these centrality measures for us. Big Ten football Twitter hashtags eigenvector centrality Big Ten football Twitter hashtags betweenness centrality As expected from the network visualization we made earlier, #B1G is the most influential hashtag in the Big Ten community. #BigTen and #Buckeyes also seem to be reasonable choices to use as a hashtag, but if we want to develop our own hashtag on Twitter, we only have room for one other hashtag. So #B1G it is! Fun fact: It's especially interesting to note here that some of the most popular hashtags in the Big Ten community, such as #Michigan, #Huskers, and #Badgers, have about the same influence in the community as the least popular hashtag, #B1GFootball. This just goes to show that even in online social networks, size isn't everything! If we go back to Hashtagify.me and compare the Big Ten community's most influential hashtag's popularity over time, we make another great finding: #B1G is the only hashtag that hasn't been tanking in popularity since the end of the primary college football season. Now we know for sure that #B1G is the hashtag we should include in our tweets if we want to get the most attention from the Big Ten community on Twitter. B1G popularity over time Learning how to use the key hashtags with Hashtagify.me Now that we've decided that #B1G is the best hashtag to use for our marketing purposes, we need to learn how to use the hashtag. After all, each online community has their own social norms and inside jokes. If we barge in with a tweet blatantly advertising our brand, we're more likely to anger the community than make them want to consider our products. That's where another handy Hashtagify.me tool comes in: hashtag top influencers. B1G top influencers @BigTenNetwork is clearly the most influential Twitter user for the #B1G hashtag, so they're likely the best user to learn the social norms from. We can take a quick scroll through their tweets to see what kind of tweet receives attention with the #B1G hashtag. Here's one successful tweet: Congrats to @MSU_Football on winning #B1GFCG and earning the conference's Rose Bowl bid. pic.twitter.com/gsr1c2NuA5 — Big Ten Network (@BigTenNetwork) December 8, 2013 Most of the successful tweets have minimal text, a large picture of something related to Big Ten sports, and are usually congratulating a sports team on their victory. There are several creative ways to construct a tweet that could market your own brand while fitting this profile, but I'll leave that to you as the brand marketer! As if all that weren't enough, Hashtagify.me takes it to the next level by even telling you when the hashtag is used the most. B1G hashtag usage We see some fairly clear trends here: #B1G is used the most on Saturdays and Sundays in the mornings before 10am EST and in the evenings after 6pm EST. Although we could employ a strategy to tweet when fewer people are tweeting on the hashtag (between 11am and 5pm EST), most likely the people who follow #B1G will only check Twitter for it during the hours that it's normally used. As such, we should tweet when #B1G is used the most to connect with more users. What are you waiting for? Stop missing out on potential customers by sending out tweets with the wrong hashtags at the wrong times. Come join Hashtagify.me today so you can maximize your brand's reach on Twitter. --- ## A Song of Ice and Fire by the numbers: Chapter titles URL: https://www.randalolson.com/2014/02/12/a-song-of-ice-and-fire-by-the-numbers-chapter-titles/ Published: 2014-02-12 Categories: data visualization Tags: a song of ice and fire, book summary, chapter titles, data visualization, game of thrones, george rr martin Randy Olson takes a bird's eye view at the "A Song of Ice and Fire" book series by looking at the chapter titles. I've been catching up on the "A Song of Ice and Fire" book series lately and ran across an Amazon book review that had an interesting take on reviewing the books. Since each chapter is named after the character it focuses on, we can get a bird's eye view of each book by counting how many chapters are dedicated to each character. Can we learn much about the books from such a high-level view? Let's take a look. Warning: possible spoilers below. Chapter titles in "A Game of Thrones" The majority of the first book is told from the perspective of the Stark family. In fact, only 25% of the book is told from another family's perspective -- and 10 of those chapters are following Daenerys on an entirely separate continent. Clearly, GRRM focused this first book on developing the Stark family, who would play a huge role in the coming crisis. Although it looks like Eddard is our manly hero, nearly half of the book is told from the perspective of women. I'm no fantasy series connoisseur, but the diversity of perspective that the book is written from really sells this book series for me. Chapter titles in "A Clash of Kings" Huh. What happened to Eddard? About 75% of the second book is still dedicated to the Stark family's perspective, but Eddard is nowhere to be found. Maybe he's taken a vacation to Volantis? With the loss of Eddard, Tyrion takes on the role as the dominant perspective in the series. Clearly Tyrion is destined for greatness in this book. Let's just hope he doesn't suddenly go on a vacation like Eddard. Lastly, we get to see the world from the perspectives of Theon and Davos in the second book, who give us a closer look at the action going on in the Iron Islands and Dragonstone. Chapter titles in "A Storm of Swords" Surprisingly, Arya takes the leading perspective in book 3. I wouldn't have expected a little tomboy to become a leading character in this series, but GRRM always has a trick up his sleeve, doesn't he? For another surprise, Samwell comes into his own in book 3, with 5 chapters dedicated to his perspective. Who would've thought a fat craven would be worth following in the frozen north? For the first time in the series, we get to see the perspective of someone who isn't a de facto "good guy": Jaime. This was an entirely new perspective for me, as most fantasy novels don't let us see through the eyes of the "bad guy." Another great selling point for the series! (Some may claim Theon fills that role in book 2, but he wasn't a "bad guy" from the beginning.) Of note: Robb never had a chapter dedicated to him, so it's strange that the TV show spent so much time on him. Perhaps it was to make the events in book 3 all that more tragic. Chapter titles in "A Feast for Crows" Cersei and Brienne burst onto the scene in book 4, two entirely new perspectives that we've never heard from before. This should make book 4 especially interesting to see what's going on in Cersei's mind. Book 4 is one of the first books that had several miscellaneous characters taking over that we haven't heard of before, and don't ever really see again (the "Other" group). Interestingly, this coincides with a whole star drop in ratings (from 4.5 to 3.5) in the ASOIAF series. I wonder if these random perspectives had something to do with it? Would GRRM have been better off cutting out the randos and focusing more on the main characters? Daenerys, Jon, and Tyrion are nowhere to be found in book 4. You bastard, GRRM! Did you send them on a vacation to Volantis too? Chapter titles in "A Dance with Dragons" Oh good! Daenerys, Jon, and Tyrion are back. I was worried there for a bit. Seeing the chapters dedicated to Davos turned out to be a bit of a spoiler for me, given the claim that he's dead in book 4. GRRM really went out of control with miscellaneous characters in book 5. I've yet to finish this book, but I hope it's not too distracting. What do you think? It looks like we can learn quite a bit about the ASOIAF series even from a bird's eye view. What do you think? Did I miss anything? Leave your thoughts in the comments below. --- ## Movies aren't actually much longer than they used to be URL: https://www.randalolson.com/2014/01/25/movies-arent-actually-much-longer-than-they-used-to-be/ Published: 2014-01-25 Categories: data visualization Tags: feature film, films, movie length, movies, visualization Randy Olson takes a look at IMDB data to see if movies are actually longer than they used to be. Every year, I hear the same complaint about movies on the big screen: Movies are getting so damn long! We're almost to the point that moviegoers should start demanding an intermission for some of these behemoths of film. Epics like The Hobbit: The Desolation of Smaug are pushing the boundaries of how long a film can be But then I started wondering: Are movies really getting longer than they used to be? Or are a few outliers--like The Lord of the Rings series--skewing our perception of what's really going on? To address this question, I turned to IMDB and gathered the 25 most popular movies from each year from 1931 through 2013. Below is the average feature film length over that time period. The blue area indicates the 95% confidence interval for feature film length each year Mean and CI have been smoothed with a rolling average (window = 5) How about that? There's several interesting stages in this data, so I'll break the analysis down by stage. 1931-1970 With the introduction of the television in the 1930s and 1940s, the movie industry suddenly had a competitor. In response, movie producers were forced to raise the bar and start producing more epic films to keep audiences packing the theater. The result? Feature films gained an extra 30 minutes between 1931 and 1960, which set the standard in film for the next 50 years, and eventually led to the blockbuster phenomenon. 1970-1985 It's strange that the average feature film lost about 10 minutes during this period. The only explanation I can think of is the videotape format war in the 1970s, where VHS and Betamax were battling it out to become the dominant movie format. Could the eventual dominance of VHS caused movie producers to keep their films shorter and well under the 2 hour mark? 1985-2000 Between 1985-2000, feature films grew back to the same length as in the 1960s. This may explain why it's usually Millennials (born 1980-2000) complaining that movies have gotten longer than they used to be: If you grew up watching movies in the 1980s, they have gotten longer for you! Meanwhile, Generation Xers are shaking their head at Millennials wondering what the heck they're talking about (as usual). 2000-2013 Perhaps the most relevant time period for us to look at is 2000-2013, because these are the movies that are the freshest in our mind. Interestingly, the average feature film hasn't gotten much longer since the turn of the century, keeping with the status quo established in 1960. This is just averages over a bunch of movies, though. What if we compare the longest feature film each year? Surely modern movies are longer than the old ones that had to fit on a VHS tape. Length of the longest feature film each year Huh. Even the maximum feature film length has hovered around 3 hours since the 1960s. It looks like movies aren't actually much longer than they used to be. We may have a few lengthy blockbusters nowadays, but they sure don't stack up to much when compared to 20th century epics like Gone with the Wind (1939, 223 mins), The Ten Commandments (1956, 220 mins), and Lawrence of Arabia (1962, 216 mins). What about looking at all films ever? Several people commented that only looking at the top 25 most popular films each year could possibly have biased this analysis, so here's the average feature film length for all feature films in the IMDB database between 1906 and 2013. The blue area indicates 1 standard deviation for feature film length each year Mean and error bars have been smoothed with a rolling average (window = 5) Although the overall average film length is much lower than the top 25's average film length, the same main trends still hold: Up until the 1950s, feature films grew by 15-30 minutes. Then after the 1950s, the average movie hovered around 90 minutes. Interestingly, the trend here shows that movies have been getting a little bit shorter in the past few years. We'll have to revisit this data in a few years to see if that trend holds. Data For those interested in the data underlying these visualizations, here's the details. The processed data is available for download here, and the raw IMDB data is available via IMDB's interface. I parsed through the IMDB "running times" list and grouped the films by year. For each year, I saved the 25 films with the most IMDB user ratings to build a list of the most popular films for each year. Number of IMDB user ratings is a reliable measure of a film's popularity because popular films that were highly successful in the box office--and thus had millions of people watching them--generally receive far more user ratings on IMDB than unsuccessful films. The above is the same reason why I picked the 25 most popular films instead of looking at all films each year: The 25 most popular films are the films that had the lion's share of people watching them in theater, thus they are a better representation of the films the average moviegoer experienced that year. --- ## A tech-focused guide to increasing your influence on Twitter URL: https://www.randalolson.com/2014/01/24/a-tech-focused-guide-to-increasing-your-influence-on-twitter/ Published: 2014-01-24 Categories: tutorial Tags: automation, branding, followers, influence, klout, twitter Randy Olson explains how you can use technology to increase your influence on Twitter without spending a dime. This is the original version of this article that appeared on Co.Labs. We all know the power of Twitter to spread one person's voice across the Internet. Celebrities tweet out pictures of their daily lives, world renowned scientists share the latest scientific breakthroughs, and even the President keeps us up to date on the politics in Washington -- all to hundreds of thousands of adoring followers, with the single click of a button. But what if you're not a celebrity, a world renowned scientist, nor the President of the United States? What if you're just an average person who wants to get your voice out there? In this guide, I'm going to go over 5 tips that are guaranteed to increase your influence on Twitter without costing you a dime. Some readers may be wondering: "Who are you to be writing a guide like this?" Well, I'm an average tech nerd who runs two of the top 5% most influential Twitter accounts in the world. I'm no celebrity by any means, but I've built these accounts up from scratch over a few months by following these tips to the letter. [caption id="attachment_2623" align="aligncenter" width="1024"] There are millions of people around the world checking into Twitter every day. It's about time you connected with them. Image c/o Eric Fischer[/caption] Understand what "influence" is on Twitter What does it mean to be "influential" on Twitter? In practical terms, it means that people actually pay attention to your tweets: They respond to your tweets, they favorite them, and if you're lucky, they like them enough to retweet them to their followers. This means that we can identify highly influential people on Twitter by finding Twitter accounts that consistently have others interacting with their tweets. It just so happens that there's a web site out there dedicated to finding those influential people on Twitter, called Klout. [caption id="attachment_2619" align="aligncenter" width="955"] The Klout dashboard menu shows your Klout score over the past 90 days. Image c/o Klout[/caption] Klout assigns a single number to each person gauging how influential they are on social networks. According to the Klout statistics, the average Twitter user -- who's lucky to get a few retweets on every other tweet -- has a Klout score of about 40. Meanwhile, the top 5% influencers on Twitter -- whose tweets are regularly retweeted dozens of times -- have a Klout score of 63 or higher. Easy enough. Of course, Klout has its naysayers. They claim that nobody cares about Klout (but you should!). They eagerly point out that Twitter bots can achieve an average Klout score. They even claim that it's ridiculous that random people on Twitter are considered more influential than real world influencers like Warren Buffet. But it's time to face the facts: Real world influence doesn't necessarily translate into influence in the Twitter world. Just how gut instinct gave way to data science in baseball in the 20th century, it's time for data science to take the throne for measuring influence on Twitter in the 21st century. If an average person -- or even a Twitter bot -- can have real users consistently interacting with their tweets, why shouldn't they be considered influential? As a side note here: Don't waste your money buying a high Klout score. If you follow these tips, you can easily join the ranks of the top Twitter influencers for free. Pick a topic for your Twitter account You won't get anywhere on Twitter if you can't pick the audience that you want to influence. Pick a topic that you care about -- or that your company cares about, if you're a marketer -- and stick to that topic. Web sites like hashtags.org and Hashtagify.me will help you find out if there's popular hashtags already established for that topic. If there aren't hashtags for your chosen topic, consider choosing a different topic: It's an uphill battle if you have to create a hashtag from the ground up. [caption id="attachment_2621" align="aligncenter" width="796"] Top 10 most related hashtags to #SEO. Image c/o Hashtagify.me[/caption] To be clear: Talking about yourself or how delicious that donut was this morning isn't the right topic. People only care about those kinds of tweets from celebrities, and let's face it, you wouldn't be reading this guide so closely if you were a celebrity. To lead, you must first learn to follow Followers are the currency of Twitter. The more people that follow you, the more likely someone's going to pay attention to and interact with your tweets. Essentially, you're playing a numbers game. If only 10% of the people who follow you are going to read your tweets, you can still get 100 people to read your tweets if you have 1,000 followers. [caption id="attachment_2622" align="aligncenter" width="511"] You must have a large following to be influential on Twitter. Image c/o Twitter[/caption] But numbers aren't everything. You can easily buy thousands of fake Twitter followers, but that won't translate into influence. You need real followers that are actually interested in your tweets, otherwise you might as well be talking to a wall. The best way to find likeminded followers is to search for tweets that use your topic's hashtags and follow the people who come up. If they don't follow you back within a few days, no worries -- just unfollow them. You can only follow so many people on Twitter, and you can't fill up valuable follow slots on users that won't follow you back. It'd be extremely laborious to do all that following and unfollowing by hand. Fortunately, there's dozens of automatic follow services for Twitter out there, and even fully loaded professional services that can follow and unfollow people in the thousands. But why pay for mass following and unfollowing when you can do it yourself for free? If you're familiar with basic scripting, there's a free Python script for DIY mass following and unfollowing. Over a few months, you can build up thousands of real followers that are actually interested in what you have to say. Acquiring real followers is the first big step toward becoming a top Twitter influencer. Use collaborative social news sites to find interesting content to tweet It's a good idea to post two or three good tweets a day so your followers come to know you as someone worth paying attention to and interacting with on a daily basis. Even the best content creators will struggle finding two or three good links to talk about every day, so how can you keep up? [caption id="attachment_2620" align="aligncenter" width="1466"] The top 10 popular links of the day on /r/Technology are a great source of articles if you want to appeal to a techie audience. Image c/o reddit[/caption] Put social news web sites like reddit and digg to work for you. Most social news web sites sort their content by category, so find the category for your topic and tweet the top links and images for that category each day. The advantage of social news sites is that thousands of people have already voted over whether they like the link or not, so you can be pretty sure that the top links will sit well with your followers. Use favorites to let potential followers know you're on Twitter The hardest part about connecting with people on Twitter is letting them know that you actually exist. With over 600,000,000 registered users on Twitter, it's easy to get lost in the crowd. If you regularly blog or closely follow some specific hashtags, go through and favorite all of the tweets that mention your blog or your favorite hashtags. Your favorite may be all the other person needs to follow you back, and now you have yet another valuable connection on Twitter. Again, manually favoriting hundreds of tweets a day isn't feasible for most people, so it's necessary to automate the process. But there's good news yet again: that same free Python script also has the option to automatically favorite all tweets that come up in a search. Give these tips a try for a couple months and watch your Twitter influence soar. Have any more tips? Leave them in the comments. --- ## More lies and bad analyses by Shareaholic are misleading the public URL: https://www.randalolson.com/2014/01/14/more-lies-and-bad-analyses-by-shareaholic-are-misleading-the-public/ Published: 2014-01-14 Categories: analysis, reddit, statistics Tags: lies, reddit, referral traffic, shareaholic, trend analysis Randy Olson debunks Shareaholic's claim that Reddit referral traffic is declining. This is the original version of the article that appeared on Co.Design. Earlier this week, Shareaholic published a report claiming that Reddit referral traffic is rapidly declining. Shortly thereafter, the report appeared on several respectable news outlets, including The Atlantic and Business Insider, chiming in and recommending that SEO experts start looking at web sites other than Reddit for referral traffic. Normally I'd ignore a report like this because I don't care much about SEO, but the visualizations and analysis in Shareaholic's report were so bad -- and so misleading -- that I had to write a response to it. It takes three points to make a trend First off, Shareaholic's claim rests on their comparison of two cherry-picked data points: Reddit referral traffic at Dec '12 and Dec '13. They compare referral traffic between those two points and note that it's gone down by 33%, then wildly speculate that Reddit's referral traffic is crashing. One of the things we learn in high school math class is that it takes at least 3 points to make a trend, not 2. In fact, if Shareaholic had posted this report in June '13, they could've told the opposite story: Reddit's referral traffic is skyrocketing (up 66%)! That's a sign right there that the methods they're using are flawed. Shareaholic's trend analysis from June 2013 Their visualization doesn't even show us if there's a trend If we want to talk about trends, we need to consider all of the data. Shareaholic shows us a simple line plot connecting the dots, but doesn't even try to show us what trend is actually going on. Can you easily tell if there's really a significant downward trend going on there? That's the sign of a poor data visualization. Reddit's traffic share according to Shareaholic The trend isn't even significant When we take a closer look at the referral traffic data Shareaholic presented, we can draw a trend line through those points and notice that there isn't much of a trend going on at all: The line is almost flat, going down only 0.007% every month. Of course, this trend line predicts that Reddit referral traffic will stay in constant decline, which we know isn't the case. Just look at what happened in June '13. There's a lot more to Reddit's referral traffic trend than meets the eye, and Shareaholic's oversimplified analysis doesn't even come close to explaining it. Shareaholic's trend analysis with all data points Even worse, if we drop the Dec '13 point entirely and plot the trend again, the line becomes even flatter (-0.004% per month). If a trend relies so heavily on one data point, then it isn't much of a trend at all. There simply isn't enough data here for Shareaholic to make any kind of claim about Reddit's traffic referral trends -- especially the audacious claims they made in the news. Shareaholic's trend analysis without December 2013 When it comes to sorting out the signal from the noise, this is pure noise that Shareaholic is trying to make look like a signal. Fractions can be misleading It's even strange that Shareaholic presented this data in terms of fraction of traffic referred from Reddit. For example, they say: "During December 2012, sites saw 0.33% of their overall traffic come from Reddit." Why not report the raw traffic numbers instead of a fraction? How Shareaholic's fractions mislead It could very well be that even though the fraction of traffic from Reddit went down, the total amount of traffic from Reddit went up. To drive this point home: Even though 50% seems less than 75%, if we're comparing 75% of 200 visitors (= 150 visitors) vs. 50% of 1,000 visitors (= 500 visitors), clearly the 50% represents more traffic overall. In other words, even if there was declining fraction of traffic from Reddit, it could simply be because the web sites Shareaholic tracks are receiving more traffic overall. With that in mind, we can't even be confident that the Shareaholic referral traffic data actually represents a decline in overall referral traffic from Reddit, as they so boldly claimed in the news. Shareaholic doesn't care if their analysis is correct Before publishing this article, I wrote to the Shareaholic team to explain the faults in their analysis. Their response? Ultimately, our goal is to provide data [that] marketers can act upon, not data that they'd sit around in a room pondering about for hours on end. It seems that Shareaholic doesn't even care if their analysis is correct, as long as it offers "data that marketers can act upon." Are you sure you want your marketers acting upon data that's as unreliable as this? If Shareaholic can't even get basic data analysis right, I don't know how much confidence that should give you in the rest of their services. --- ## A look at sex, drugs, violence, and cursing in film over time through MPAA ratings URL: https://www.randalolson.com/2014/01/12/a-look-at-sex-drugs-violence-and-cursing-in-film-over-time-through-mpaa-ratings/ Published: 2014-01-12 Categories: data visualization Tags: cursing, drugs, films, language, movies, MPAA rating, sex, violence Randy Olson explores sex, drugs, violence, and cursing in films over time by looking at MPAA film ratings. In 1968, the Motion Picture Association of America (MPAA) film rating system took effect in an effort to replace the then-aged Hays Code. The goal of the MPAA rating system is to provide a modern, standardized scale on which to rate the suitability of films for certain audiences in the U.S.A. and its territories. As such, the MPAA rating system judges every film and assigns them to one of five categories: G (suitable for everyone), PG (parental guidance suggested), PG-13 (parents strongly cautioned), R (for adults), and NC-17 (no one under 18 allowed to watch) These categories have changed a little over the years, but these are the ratings as we know them now. Although the MPAA film rating system is voluntary, films aren't allowed to be shown in most theaters nowadays without a MPAA rating. Therefore, MPAA ratings provide a useful snapshot of how major films in the U.S.A. have changed over the years. Starting in late 1990, the MPAA started including reasons for the ratings they assigned, including "sexual content", "drug use", "violence", and "strong language". I tallied these reasons for every year and ratings category and plotted them below. These tallies provide an interesting view of sex, drugs, violence, and cursing in film over time from 1990 through 2013. Overall trends in MPAA ratings reasons Reasons for PG or higher MPAA Rating The shocker here is that by 2013, 5 out of every 10 big screen films with a PG rating or higher has some form of sex and/or violence in it, and 8 out of every 10 films have some form of cursing. I've become so desensitized to sex, violence, and cursing in films that I barely even realize it, but even mainstream comedies like 21 Jump Street (2012) are rated R for sex, drugs, violence, and cursing! Wow! The sharp increase of sex in 1990s films was no surprise for me: One of the biggest changes I remember in 1990s films was the sudden burst of sex scenes in every popular film. Even in films where it didn't make sense for there to be a sex scene, suddenly I'd see the hero and heroine going at it like monkeys for several minutes while I awkwardly shifted in my seat next to my parents. Another, more subtle trend we see here is a slight increase of drug usage in films in the 1990s. This could be seen as a reflection of the U.S.A.'s slowly relaxing attitude toward alcohol and marijuana usage over time. It's too bad we don't have any data on drug usage in films in the 1960s and 1970s! PG rated films Reasons for PG MPAA Rating This is one of the most comforting graphs that came out of this project because it shows us that at least PG rated films -- which most kids and teenagers are able to watch nowadays -- don't have drugs nor sex in them (for the most part). Another trend we see over time is that violence and cursing are slowly going down in PG films, which should make us feel a little more confident in letting our kids watch newer PG rated films. This graph is also one of the funniest graphs from this project because it shows how strange U.S. culture is about sex, drugs, and violence. Especially in the 1990s, if there was violence or cursing in a film, it was usually okay for kids to watch. But if there's any hint of sex or drugs in the film? No way, Jose! We can't expose our kids to that! PG-13 rated films Reasons for PG-13 MPAA Rating The PG-13 ratings reasons seem to be a wash. Other than the trend of increasing drug usage in PG-13 films, there isn't much to see here trend-wise. R rated films Reasons for R MPAA Rating Here's where things get interesting again. We can tell that the biggest rise of sex-related scenes in the 1990s happened in R-rated films. Although it's always difficult to attribute causation to a correlation, it seems pretty curious that less than a year after the MPAA started including reasons for their ratings, R-rated films with sex scenes suddenly spiked. Could it be that directors started cramming sex scenes into their films in the 1990s to get an R rating so their films would be more appealing to adults? As they say: Sex Sells, and what's a better way to advertise that there's sex in your film without being obvious about it? We're starting to see a similar trend with drug usage in films, too. Now that drug usage is becoming ever more hip, more and more films every year are including drugs in an attempt to connect with moviegoers. How long will it be until it becomes commonplace for every R-rated, adult film to have sex, drugs, violence, and cursing? Data For those interested in the data underlying these visualizations, here's the details. The processed data is available for download here, and the raw IMDB data is available via IMDB's interface. I parsed through the IMDB "MPAA ratings reasons" list and looked for the phrases "sex", "drug", "violence", or "language". This of course misses a few films where they didn't mention any of those words in the ratings reasons -- Braveheart being the most notable example that I noticed -- but it still captures most of the 12,000+ film's ratings correctly. I excluded NC-17 in these graphs because there were so few films (< 4 per year) that were rated NC-17. What's that say about the commercial viability of the NC-17 rating, eh? --- ## 3 easy steps to avoid deceitful data visualizations URL: https://www.randalolson.com/2014/01/06/3-easy-steps-to-avoid-deceitful-data-visualizations/ Published: 2014-01-06 Categories: data visualization, tutorial Tags: best practices, data visualization, deceitful, infographic, lie Randy Olson explains how you can save yourself from getting duped by bad infographics. This is the original version of the article that appeared on Co.Design. We live in an age of Big: Big Computers, Big Data, and Big Lies. Faced with an unprecedented torrent of information, data scientists have turned to the visual arts to make sense of big data. The result of this unlikely marriage -- often called "data visualizations" or "infographics" -- have repeatedly provided us with new and insightful perspectives on the world around us. However, time and time again we have seen that data visualizations can easily be manipulated to lie. By misrepresenting, altering, or faking the data they visualize, data scientists can twist public opinion to their benefit and even profit at our expense. Due to our natural tendency to trust images more than text, we’re more likely than ever to be fooled by data visualizations. Fortunately, there are 3 easy steps we can follow to save ourselves from getting duped in the data deluge. Visualization via visual.ly Check the data source Not all data are created equal. Here's how to sort out the good data from the bad. Make sure the data source is reliable Data collected by an amateur is more error-prone than data collected by a professional scientist. Do a quick web search to see if the people who collected and organized the data have a good track record of collecting and distributing data. Make sure the data source isn't biased A drug company may be inclined to present fake data showing that their latest drug is more effective than it really is, or a political organization may manipulate data to discredit their political opponents. Think twice when considering data provided by biased groups. Generally, we can trust data provided by: Government organizations University research laboratories Nonpartisan organizations And we should look more closely at data provided by: For-profit companies Partisan organizations Advocacy groups If the data source isn't listed, take the data visualization with many grains of salt. Check the data alterations All data sets require a little bit of house cleaning before they can be visualized, but excessive curation can be a sign of misrepresented data. Every good data visualization will come with the blueprints describing how the data was manipulated from its raw form into the visualization you see. Give the blueprints a quick read and watch out for the following data alterations. Excluded data Ensure that the explanations for excluding that data are reasonable. Sometimes the "explanation" may be that the data inconveniently contrasted with the story the author wanted to tell. Transformed data Data transformations can complicate the relationships between data. It's difficult to interpret a finding such as, "The log transform of a city's productivity is related to the log transform of the city's population." Be wary if several transformations have been applied to the data. Statistics Statistics are an often-abused tool in data science. "Fatal shark attacks have risen 100% this year" sounds like an alarming statistic until you realize that only 1 person was fatally attacked by a shark last year. Check the raw numbers when data visualizations present only the statistics. Comparing statistics is even trickier. If a survey shows that 50% of Latinos and only 30% of Caucasians enjoy watching baseball, those results could easily have been purely due to chance because the survey only interviewed 20 people of each ethnicity. If the visualization doesn't indicate their confidence in the comparison (called statistical significance), then we shouldn't be confident in their comparison. If the details on the data alterations aren't provided with the visualization, always keep in mind how easy it is to make data lie when it's visualized. Check the data presentation The subtlest way a data visualization can fool you is by using visual cues to make data stand out that normally wouldn't. Be on the lookout for these visual tricks. Color cues Color is one popular tool for making certain data stand out above the rest. When considering the map below, Kentucky and Utah (the darkest and the lightest) will most likely stand out to us first. Map visualization c/o IBM Many Eyes If this map were showing percentage of the population that smokes (dark = more, light = fewer), we may quickly conclude that Kentucky has a serious smoking problem. But what if we looked at the raw numbers and saw that 27% of Kentuckians and 23% of Utahans smoke? Not so big of a difference after all. Make sure to look at what the colors actually represent before drawing a conclusion from the visualization. Structural cues Structure is another popular tool for making data immediately stand out. In the bar charts below, we're looking at the same data but with different y-axis ranges. Notice how such a simple structural change can make data look much more significant. Is an increase of 15 fraudulent visualizations from last year really "skyrocketing"? Don't let the structure of the visualization decide that for you. Always check the numbers that the visualization is representing. Remember: To save yourself from getting tricked by deceitful data, check the data source, alterations, and presentation. If you liked what you saw in this post and want to learn more, check out my Python data visualization video course that I made in collaboration with O'Reilly. In just one hour, I will cover these topics and much more, which will provide you with a strong starting point for your career in data visualization. --- ## Top 25 most murderous directors of all time URL: https://www.randalolson.com/2014/01/05/top-25-most-murderous-directors-of-all-time/ Published: 2014-01-05 Categories: data visualization Tags: directors, kill counts, movies, murderous, violence, violent films Randy Olson visualizes and discusses the top 25 deadliest actors of all time. Last weekend, I discovered a web site where a group of dedicated film fanatics have been systematically counting on-screen deaths in various films for several years. I immediately set out to scrape all of the data from their web site and forums so I could visualize it and see what we could learn. I posted the data online to save everyone else the chore of repeating the scrape. Below is the fourth in a series of visualizations I created, showing the top 25 most murderous directors ordered by the number of on-screen deaths in their films. If you would like to feature any of these visualizations on your own web site, please contact me first. Top 25 most murderous directors of all time Some interesting facts The most murderous female director? Katja von Garnier with 256 on-screen deaths in films she's directed. Most of Peter Jackson's film deaths happen in just two films. 836 in LotR: Return of the King (2003) and 468 in LotR: The Two Towers (2002). This picture really helps Peter Jackson's case for being the most murderous director of all time: Peter Jackson, the most murderous director of all time John Woo seems to have stopped actively directing films in 2003, so even though Jackson and Woo are ranked closely, it looks like Jackson will easily maintain his lead -- especially with The Hobbit films coming out! Sylvester Stallone ranked in 3rd on the deadliest actor list and 15th on the most murderous directors list. Unsurprisingly for some, Stallone has truly built a career off of violence and death on film! Have you found any interesting facts in the data? Post them as a comment here! --- ## Has violence been vanquished? Not in the film industry URL: https://www.randalolson.com/2014/01/04/has-violence-been-vanquished-not-in-the-film-industry-2/ Published: 2014-01-04 Categories: data visualization Tags: deadliest actors, film, kill counts, steven pinker, violence, worldwide Randy Olson discusses violence in the film industry and its implications for modern society. Since I posted my deadliest actors and deadliest films lists last week, I've been bombarded with questions about violence in the film industry. "Has the film industry become more violent over time?" (My answer: yes!) "Why do you think the film industry portrays so much violence?" "Is it sad or depressing that the film industry portrays so much violence?" (scroll down to the bottom) Even a renowned public news station recorded an interview with me this week to talk about it on the radio. Unfortunately, the producers decided against running the story, but that interview convinced me that I need to write my thoughts on the subject down somewhere. So, what better to do with a dropped interview than turn it into a blog post? If this sounds like a news story you'd like to run, send me an email and let's set up an interview. What motivated you to create the deadliest actor list? Short answer: I did it for fun. I'm an AI researcher at Michigan State University, and every weekend I've been practicing my data visualization skills. Instead of playing video games or going to the movie theater, I scour the internet for an interesting data set and try to visualize it in a way that makes us look at the world a little differently -- and maybe even understand it a little better. On the weekend I made this visualization, I ran across MovieBodyCounts.com, which has a community of several dedicated film fanatics who have been systematically counting body and kill counts in films for over 5 years. I found their work incredibly interesting, so I decided to bring their work to life in a series of data visualizations. Is there anything surprising to you about the deadliest actor list? I was surprised by the utter lack of John Wayne. John is well-known for killing hordes of bad guys in his pre-1960s Western flicks, but he doesn't show up anywhere on this list. Sadly, John's movie kill counts are missing from the kill count database. I've been trying to recruit film-watching heroes to do those counts so we can see how John Wayne stacks up. John Wayne never needed to kill more than 5 people in a movie Any fun facts about the deadliest actor list? Uma Thurman ranks in as the deadliest woman with 77 on-screen kills. Arnold Schwarzenegger's highest single-film kill count comes from Commando (1985), where in the final island scene he racked up 74 kills: 2 throats slit, 51 people shot, 1 person stabbed, 2 people stabbed by circular blades, 5 people blown up by grenades, 5 people blown up by rocket launcher, 7 people blown up by planted explosives, and 1 unfortunate person impaled. The highest kill count by any one person in a film was attained by Tomisaburo Wakayama, who scored a shocking 150 kills in Lone Wolf and Cub: White Heaven in Hell (1974). My favorite fact? While preparing for the interview, I found out that Arnold Schwarzenegger is trying to run for President in 2016. Does that remind you of a particular movie with a bad ass, machine gun-wielding action hero as President? President Camacho from Idiocracy Is it sad or depressing that the film industry portrays so much violence? What does this say about modern society? Societies have always had some form of violent entertainment. In the days of the Roman Empire, Gladiator battles -- where real people fought and died -- were extremely popular. In the Medieval age, there were tournaments where knights battled each other for glory and entertainment. Even during the Renaissance -- touted as one of the most enlightened periods of European history -- some of the most famous plays (e.g., Hamlet) portrayed violent scenes where several people were killed. It's actually an improvement that society has moved from real people dying for entertainment to entirely fake deaths on film. If you follow Steven Pinker, you've likely heard about his recent work showing that violence has been vanquished: "We believe our world is riddled with terror and war, but we may be living in the most peaceable era in human existence." Given the data below, it's hard to argue against Pinker. Worldwide violence has been on the decline Could one consequence of declining real world violence be that we turn to artificial violence -- such as film and video games -- to satisfy our innate need for violent entertainment? Possibly. I've even found that violence in film has been on the rise at the same time that real-world violence has been declining. All things considered? I don't think it's depressing at all that our film industry portrays so much violence. --- ## Top 25 deadliest actors of all time by on-screen kills in movies URL: https://www.randalolson.com/2013/12/31/deadliest-actors-of-all-time-by-on-screen-kills-in-movies/ Published: 2013-12-31 Categories: data visualization Tags: actors, deadly, kill counts, movies, violent films Randy Olson visualizes and discusses the top 25 deadliest actors of all time. Over the weekend, I discovered a web site where a group of dedicated film fanatics have been systematically counting on-screen kills for various actors for several years. I immediately set out to scrape all of the data from their web site and forums so I could visualize it and see what we could learn. I posted the data online to save everyone else the chore of repeating the scrape. Below is the second in a series of visualizations I created, showing the top 25 deadliest actors by on-screen kill counts. If you would like to feature any of these visualizations on your own web site, please contact me first. Top 25 deadliest actors of all time by on-screen kills in movies Some interesting facts Uma Thurman ranks in as the deadliest woman with 77 on-screen kills. Arnold Schwarzenegger's highest single-film kill count comes from Commando (1985), where in the final island scene he racked up 74 kills: 2 throats slit, 51 people shot, 1 person stabbed, 2 people stabbed by circular blades, 5 people blown up by grenades, 5 people blown up by rocket launcher, 7 people blown up by planted explosives, and 1 unfortunate person impaled. The highest kill count by any one person in a film was attained by Tomisaburo Wakayama, who scored a shocking 150 kills in Lone Wolf and Cub: White Heaven in Hell (1974). Have you found any interesting facts in the data? Post them as a comment here! --- ## Top 25 deadliest films of all time by on-screen death counts URL: https://www.randalolson.com/2013/12/31/deadliest-films-of-all-time-by-on-screen-death-counts/ Published: 2013-12-31 Categories: data visualization Tags: deadly, films, kill counts, movies, violent films Randy Olson visualizes and discusses the top 25 deadliest films of all time. Over the weekend, I discovered a web site where a group of dedicated film fanatics have been systematically counting on-screen deaths in various films for several years. I immediately set out to scrape all of the data from their web site and forums so I could visualize it and see what we could learn. I posted the data online to save everyone else the chore of repeating the scrape. Below is the first in a series of visualizations I created, showing the top 25 deadliest films by on-screen death counts. If you would like to feature any of these visualizations on your own web site, please contact me first. 25 deadliest films by on-screen death counts I also created a Top 100 list if 25 just isn't enough. Some interesting facts Despite being rated PG-13, Lord of the Rings: Return of the King has the most on-camera deaths in a movie in all time. Across all of the Lord of the Rings movies, Legolas in fact won the on-screen kill count contest with Gimli, with 57 on-screen kills vs. 25 on-screen kills for Gimli. Despite its claim to fame as being the bloodiest film of all time, Hot Shots Part Deux ranked in with only 114 on-screen kills. Orlando Bloom is in 4 of the top 8 deadliest films of all time. Have you found another fun fact that you'd like to share? Post it here in the comments. FAQ This visualization has made its way around the internet. Here's some common questions I heard asked about it. What about movies where a building/airplane/planet/death star exploded? That counts for hundreds/thousands/millions of on-screen deaths. The web site I took the data from only counts visible, on-screen deaths. Implied deaths, such as an entire planet exploding, are not counted. A movie I know of is missing that definitely has at least (large number) of deaths in it. The folks gathering the data have counted most of the movies with high kill counts. Although they have been at it for several years, this is by no means a comprehensive list. If you're willing to do a body count for the movie, awesome! Check their kill counting guidelines and post a report on their forums. I don't believe the body count for one of the movies. You can go to the data source's Movie Index and double-check the counts for yourself. They provide a scene-by-scene breakdown. If 600 people died in the movie 300, does that mean the Spartans had a 1:1 killing ratio? No. The movie did not show all 300 Spartans dying, and there were several implied deaths on the Persian's side. --- ## Top 25 most violence packed films of all time URL: https://www.randalolson.com/2013/12/31/most-violence-packed-films/ Published: 2013-12-31 Categories: data visualization Tags: deadly, films, kill counts, movies, violence, violent films Randy Olson visualizes and discusses the top 25 most violence packed films of all time. Over the weekend, I discovered a web site where a group of dedicated film fanatics have been systematically counting on-screen deaths in various films for several years. I immediately set out to scrape all of the data from their web site and forums so I could visualize it and see what we could learn. I posted the data online to save everyone else the chore of repeating the scrape. Below is the third in a series of visualizations I created, showing the top 25 most violence packed films ordered by on-screen deaths per minute. If you would like to feature any of these visualizations on your own web site, please contact me first. Top 25 most violence packed films of all time I also created a Top 100 list if 25 just isn't enough. Some interesting facts Only 39 Spartans were shown dead or dying in the movie "300." That means "300" shows 4.79 dead or dying Persians per minute. Even if the Spartans kept up the killing at that rate 24/7, they would've taken at least 2 weeks to wipe out the entire Persian army. Most of the on-screen deaths in Lord of the Rings: Return of the King happen in the relatively short Battle of the Pelennor Fields: Rohirim killed – 123 Haradrim killed – 32 Denethor - 1 Nazgul killed – 1 Orcs killed - 239 Orcs ridden down by Rohirim – 83 Rambo (2008) started the carnage as early as possible with several villagers being gunned down only 3 minutes into the movie. TMNT (2007) is the only PG-rated movie that ranks in this top 25 list. Have you found another fun fact that you'd like to share? Post it here in the comments. --- ## Violence has been on the rise in the film industry URL: https://www.randalolson.com/2013/12/31/violence-has-been-on-the-rise-in-the-film-industry/ Published: 2013-12-31 Categories: data visualization Tags: film, kill counts, movies, rising violence Is the film industry becoming more violent over time? Randy Olson explores this question by looking at the data. When I shared my earlier visualizations showing the deadliest films and actors of all time, several people commented that there appears to be a trend of increasing on-screen deaths over time. To follow up on that observation, I analyzed the film death count data to see if that trend actually holds. In the visualization below, I grouped all of the films by year and summed all of the on-screen death counts for each year. I also plotted the Pearson's correlation coefficient to give a sense of how total on-screen body count and year are correlated. Rising Violence in Films Just by looking at the black trend line, it's clear that the total death count has been on the rise, especially after 2000. If the raw data alone doesn't convince you, the statistics will: A correlation coefficient of +0.71 indicates that year and total death count are highly correlated! In other words: A film will generally have more on-screen deaths the later it is filmed. What are the causes of this increase in violence? Now that we have clear data showing a trend of increasing violence in the film industry, we are left to wonder: What are the causes? I don't think it's because our culture has become more violent. Several studies have shown that the world is becoming less violent as time goes on. Could it be that instead of acting on our violent tendencies, we're sating them with violent films, video games, sports, etc.? I don't have the data on video games, but I suspect a similar trend would hold. Sports have always been popular, but have they become more violent over time? The rise in popularity of MMA and similar violent sports seems to say yes. Another, perhaps more pragmatic, consideration is that advances in the film production industry have enabled movie producers to show massive, epic battles (and the death tolls that come with them!) that they've always wanted to show. If a pre-2000 director wanted to show an epic battle in a film, he/she would have to hire hundreds/thousands of actors to act out the battles in real-time. (Which isn't unheard of; props to Waterloo and Braveheart!) In contrast, a relatively small team of animators today can create a breathtaking landscape and massive armies with a couple weeks of computer animation. Just look at Lord of the Rings: Return of the King's Battle of the Pelennor Fields scene. There's no way a scene like that could happen without computer animation. What do you think? Is the rise of violence in the film industry simply because it's easier to kill animations than actors, or is there a deeper reason? Post your thoughts in the comments here. --- ## Political polarization of the U.S. Senate: Is the data fooling us? URL: https://www.randalolson.com/2013/12/21/political-polarization-of-the-senate-is-the-data-fooling-us/ Published: 2013-12-21 Categories: analysis, data visualization Tags: data visualization, network analysis, U.S. Senate, voting behavior Randy Olson suggests that the data visualizations showing the sudden "polarization of the U.S. Senate" may be telling us a lie. Earlier this month, The Economist, Yahoo! News, and several other respectable news outlets ran articles talking about some great network visualizations apparently showing the "political polarization of the U.S. Senate." There, they argued that these visualizations show how the Senate has evolved from a fairly cohesive unit in 1989 into a dysfunctional group divided along party lines in 2013. If you look at the visualizations of voting behavior in the U.S. Senate below, you'd probably agree with their conclusions. The supposed "political polarization" of the U.S. Senate Visualization c/o Renzo Lucioni I was a little skeptical. There were a couple issues with how the visualizations were presented: The divide over time wasn't as clear as the news articles made it appear, which made me suspect that they were cherry picking. All connections between Senators who voted on less than 100 bills together were arbitrarily removed, which could have produced the sudden "polarization" effect we were seeing. Renzo was kind enough to point me to his script that collected all the data for these visualizations, so it was easy enough for me to collect the data myself and run my own analyses. I've provided the data online free to download on figshare. Were the news outlets cherry picking to prove their point? To address the first concern, I measured the modularity of the networks over time. This modularity score basically gives us a measure of how much the Senators are divided up into disconnected groups, or political parties in this case. In the graph below, I called this measure "divisiveness." Divisiveness of the U.S. Senate, quantified The x-axis is time, and the y-axis is the modularity score It's quite amusing to see that the first two network visualizations that The Economist showed -- that were meant to show the Senate as a fairly united unit -- were at the two points of lowest divisiveness in the entire time period (1989 and 2002). Similarly, the final network visualization that they showed -- that was meant to show a divided Senate -- was at a point when divisiveness was second highest (2013). Coincidence? For the choice of the 1989 and 2002 visualizations, I don't think so. It's fairly clear that the news outlets were cherry picking the visualizations to prove their point. Does the data really show a divided Senate? My second concern was that, by removing all connections between Senators who voted on fewer than 100 bills together, Renzo could have produced the sudden appearance of "divisiveness" in the Senate when the Senate has always been divided. If you look at the "divisiveness, quantified" graph above, the Senate has always been fairly divided since the early 1990's. If that's the case, why do Renzo's visualizations look fairly cohesive early on, but suddenly divided around 2013? This interactive visualization on a Yahoo! News article illustrates my concern best: If you cut many connections, the Senators form into clusters based on political affiliation. If you don't cut any connections, then the divide is much less clear. Cutting many connections produces a divided Senate in 2013 Screen shot taken from this visualization Cutting fewer connections produces a much less divided Senate in 2013 Screen shot taken from this visualization Moving the cutoff threshold around easily changes whether it looks like the Senate is divided or not. In fact, if you play around with that interactive visualization enough, you'll see that you can eventually produce a politically divided Senate for every year by removing enough connections. Or, if you want, you can produce a politically united Senate by leaving more connections. So what does this tell us? Politicians have always been more likely to vote along party lines, even in 1989. Perhaps they've been slightly more inclined to do so this year than in 1989, but it's not as extreme as Renzo's visualizations suggest. The political parties only appear fairly cohesive in 1989 and completely divided in 2013 because of an arbitrary cutoff, and we could tell a completely different story if we chose a different arbitrary cutoff. Always be skeptical of data visualizations The point of this post wasn't to rail against Renzo nor any of the news outlets that hyped his network visualizations. Instead, I hope you found this to be a cautionary tale to always question how data visualizations are made. Data visualizations can often be manipulated to demonstrate any point the author pleases, and it's easy to accept what the visualization claims when it agrees with our intuitions. Mark Twain once famously wrote: "There are three kinds of lies: lies, damned lies, and statistics." I'd like to expand on his quote by adding one more kind of lie: "There are four kinds of lies: lies, damned lies, statistics, and data visualizations." --- ## Sharing your passion for science with the world through reddit: Interview with Unidan URL: https://www.randalolson.com/2013/11/14/sharing-your-passion-for-science-with-the-world-through-reddit-interview-with-unidan/ Published: 2013-11-14 Categories: outreach, reddit Tags: Ben Eisenkop, reddit, science outreach, social media, unidan Randy Olson interviews "reddit celebrity" Unidan about his experience teaching science to the public on reddit. For the third and final interview in this series of posts about science outreach on reddit, I'm interviewing a "reddit celebrity" who became famous for sharing his passion for science with the rest of the world through reddit. If you missed the previous interviews, I've already interviewed a postdoc who gave a reddit IAmA and a professor who runs reddit's AskScience forum. Hopefully I've convinced you by now that it's more than worth your time to engage in science outreach through reddit, but if not, give this post a read and see if it floats your boat. Today I'm interviewing reddit user Unidan, affectionately known to reddit users as "The Excited Biologist." Outside of reddit, Unidan is known as Ben Eisenkop, and works at a public university as a graduate instructor. When he isn't teaching students at his university or imparting random bits of knowledge to the public on reddit, Ben works as an ecosystem ecologist, primarily working with wild bird populations. [caption width="400" align="aligncenter"]Ben Eisenkop, a.k.a. Unidan, is always excited about science Photo c/o Ben Eisenkop[/caption] "Biologist here!" Ben became famous on reddit when he started showing up in random conversation threads and imparting tidbits of scientific knowledge related to the conversation. He'd always start the comment with the phrase, "Biologist here!," then proceed to go into excessive detail about the most random facts. Here are a few examples: [1] [2] [3]. One of Unidan's early comments on Reddit Ben's enthusiasm and scholarly tone won over the hearts and minds of reddit's users, who demanded that he run an IAmA session so they could ask him everything they want to know about science. That IAmA ended up becoming one of the most successful IAmAs on reddit to date and ran for over 5 months, reaching millions of users and inspiring them to learn more about science. Whenever someone has a random science question on reddit, inevitably someone will begin chanting "/u/Unidan, /u/Unidan, /u/Unidan!" in hopes of summoning him to the thread to share his wisdom. In return for his time and knowledge, reddit users shower him with gifts and adulation. There's even talk of Ben running a nature-focused YouTube series backed by reddit. Doesn't this sound like the kind of interaction you'd love to have when engaging with the public? Below is my interview with Ben. My questions are bolded. Interview with Ben Eisenkop (Unidan) What motivated you to use reddit as a science outreach tool? Well, I started using Reddit more as a time-wasting device, actually, like most people do! It wasn't until I started releasing little tidbits from my day or things that I've accumulated in my work that people began to really take interest. Eventually, I was asked to do an AMA which took off while I was asleep, and I woke up to quite a myriad of questions! How do you use reddit to communicate science to the public? I try to be very conversational. In my opinion, Reddit's semi-anonymity is perfect for straining out bad questions, but also letting people ask incredibly honest ones that even my real students would be very reluctant to ask, due to how they may be perceived by others. For me, that lets me answer some very fun questions and look at problems from angles that I've never considered before, especially when I get to speak to people outside of the biological sciences. What are the pros and cons of using reddit as a science outreach tool? The pros are that I can really respond at my leisure, and I can gather myself up pretty nicely and revise exactly what I want to say and convey instead of having to do it immediately on the spot. I respond, in my own opinion, pretty quickly, regardless, but it's still nice to be able to fact-check myself before I give someone advice, and it's also great because I can converse with people for months or sometimes years in sparse little pieces all over the globe! It's been incredible to talk to people with such a wide variety of experiences and backgrounds. What was your best experience engaging with the public on reddit? Have you had any negative experiences while engaging with the public on reddit? My AMA was extremely positive, and went on for six months straight! I'd say most people that I interact with are incredibly positive, and I've even managed to change the minds of a few of the bad ones. People's opinions of me are certainly normally distributed though, as there's always people on the extreme negative side, too: I've gotten death threats, things sent to my actual address, etc. I don't worry too much, though, I figure most of that type of stuff are just pranks gone a bit too far! :D What other science outreach programs have you participated in? How would you compare science outreach on reddit to other science outreach programs you have participated in? I've done talks at schools from time to time, and symposium visits within academia, of course. I was recently invited to go to SUNY ESF and Syracuse University and was able to convince a colleague of mine to help me bring a falcon with us and, together, we gave a talk on the ecology of birds, which was a lot of fun! Here's a video of the falcon from the talk that day. Comparing the experiences is a big difference, people commend me for answering a wide variety of questions, but honestly I think answering extremely technical questions from members in your own field is much more difficult. I've gotten some really intense questions from other biogeochemists in person, and those are the real thinkers, instead of the more trivial type of questions I'm often asked on Reddit. At least online, I'm not often asked to solve difficult chemistry equations! What was it like giving a reddit AMA as a scientist? It was a ton of fun, I had a blast! Near the end, I felt bad because I was often directing people to incredibly long diatribes I had already written as many of the questions were beginning to repeat themselves, but there was a certain sense of accomplishment that came with helping a lot of younger people find direction and solve specific problems. One of the funniest things was the massive outpouring of homework questions that came via PMs on Reddit, I've never done private tutoring, but now I feel like I can list it on a resume, haha! Do you have any advice for scientists interested in engaging with the public through reddit? Never talk down to people, always give them the benefit of the doubt and talk to the height of your intelligence. That's not to say be incredibly verbose or overly didactic (like using the words "verbose" and "didactic"), but rather, assume your audience is able to follow along. If people feel like they're being talked down to by some "knowledgeable person," they'll tune out, regardless of the message. And if you're wrong about something: admit it! It happens to me all the time. I'm probably wrong about something right now! It's nothing personal, and admitting it doesn't make you less of anything, in my opinion, if anything, it makes you much more credible. Want to share your knowledge with the world? It takes less than a minute to sign up for reddit. Conversations about science happen every day on reddit, and are perfect opportunities to practice how you communicate science and share your knowledge with the world. What are you waiting for? Give it a shot! --- ## The curious case of the closed access data set in the open access journal URL: https://www.randalolson.com/2013/11/13/the-curious-case-of-the-closed-access-data-set-in-the-open-access-journal/ Published: 2013-11-13 Categories: open science, philosophy Tags: data sharing, open access, open science, plos one, proprietary data, twitter Randy Olson ponders how studies using proprietary data should be handled when published in open access journals. Earlier this year, I ran across a news article that got me really excited in the science-nerdy kind of way. The article talked about how we could measure how happy the people in each U.S. state are just by looking at geotagged tweets. They even linked to a shwanky web app that the researchers had put together showing the "average happiness" of the U.S. since 2009. I have a penchant for playing around with social network data, so I was ecstatic when I saw that the authors had published the corresponding article in PLoS ONE (two articles, in fact). That means I could easily get my hands on the raw data set, right? Wrong. I received a response to my raw data request a week later saying that they couldn't share the raw Twitter data. They're absolutely right, of course. It says it right there in the Twitter API usage terms. Twitter has made it clear time and again that they don't want Twitter content being stored outside of Twitter, and they especially don't want people sharing that Twitter data if it is stored external to Twitter. Basically: If you want Twitter data, you have to go to them and access it through the Twitter API yourself. Here's a couple clauses in the Twitter API usage terms that make it difficult to use for research: You shall not use Twitter Content or other data collected from end users to create or maintain a separate status update or social network database or service. You will not attempt or encourage others to use or access the Twitter API to aggregate, cache (except as part of a Tweet), or store place and other geographic location information contained in Twitter Content. The problem is: That restriction directly contradicts PLoS ONE's rules about sharing data. In fact, it's bolded right there on the web site: PLoS ONE will not consider a study if the conclusions depend solely on the analysis of proprietary data. PLoS ONE's stance on proprietary data makes sense. After all, one of the major reasons PLoS was founded was to make research easily accessible and reproducible -- and that entails sharing the raw data underlying every study. So, what can be done about this curious case of the closed access data set in the open access journal? Does this mean researchers using Twitter data can't publish in open access journals? Does this render Twitter an inviable platform to study social networks, if the ultimate goal is to publish the study open access? I don't really have any answers, and the folks at PLoS ONE have been pondering it since June. Any thoughts? Update (11/17/2013) -- possible solution? After a brief email conversation with Jonathan Eisen (partially shown here in the comments), we reached a couple possible solutions: Jonathan Eisen ... I think the only way I would ponder allowing something to be published would be if the full workflow for ALL analyses of said data was released so that at least people could examine and try to use the workflow themselves. If not, I don't like it. Randy Olson The authors explained their method in the paper. However, a critical component of replicating the study is accessing Twitter to get the tweets from 2011 that they actually used for the study, which by now is extremely difficult if not impossible. (By default, the Twitter API only accesses recent tweets.) IMO, that makes the study irreproducible. Some possible solutions given Twitter's data sharing restrictions: 1) If researchers could denote which tweets they used in a study (e.g. with a list of tweet IDs) and Twitter allowed the mining of specific tweet IDs, then the study is semi-reproducible. The person replicating the experiment still has to mine all 10,000,000 tweets, which is a significant burden, but at least it's possible to access the same data used in the study again. 2) If Twitter could allow researchers to register a set of tweets in Twitter with a key name (e.g., "geo-happiness-plos-one-2013-tweets"), then researchers reproducing the study could contact Twitter and ask for the set of tweets with that key name. That of course places a burden on Twitter to organize tweets in a certain way, which I doubt they will do (unless there's $$$ in it). What do you think? --- ## Answering people's pressing science questions on reddit: Interview with Tobias Landberg URL: https://www.randalolson.com/2013/11/12/answering-peoples-pressing-science-questions-on-reddit-interview-with-tobias-landberg/ Published: 2013-11-12 Categories: outreach, reddit Tags: askscience, reddit, science outreach, social media Randy Olson interviews /r/AskScience moderator Tobias Landberg about his experience teaching the public about science on reddit. This post is the second in a series of posts where I am interviewing scientists who do science outreach to the public on reddit. My goal here is to discuss these scientist's experiences to give everyone a taste of what science outreach is like on reddit. If you're still on the fence about whether reddit is a good place to do science outreach, give these interviews a read and see if it suits your fancy. My second interview is with Tobias Landberg, an Assistant Professor of Biology at Arcadia University. Tobias holds an undergraduate and master's degree in Organismal Biology from the University of Massachusetts Amherst, and a Ph.D. from the University of Connecticut. Prior to becoming a professor, Tobias worked at two postdoctoral positions: One at Boston University / the Smithsonian Tropical Research Institute in Panama, and another at the Watershed Studies Institute at Murray State University. Tobias has been moderating and answering questions on reddit's AskScience forum for over three years now, drawing on his broad research experiences to educate the public about science. Tobin Hieronymus and Tobias Landberg in their natural habitat. /r/AskScience AskScience is one of reddit's default forums, meaning that any new reddit user automatically subscribes to the forum. Currently, there are over 1.3 million users that subscribe to and read posts on AskScience on any given day. AskScience is unlike any of the other default reddit forums because it's dedicated solely to asking science-related questions. The /r/askscience front page Users vote the questions up or down depending on how good they think the question are. Good questions float to the top of the page, while inane questions sink to the bottom. This means that the community conveniently sorts out the questions that they are most eager to hear the answers for. A typical question on /r/askscience In any given thread, scientists and AskScience panelists offer answers to the question and link the readers to scientific articles if they want to read more. Just like with the questions, users also vote up or down on the answers, so the best answers usually float to the top. With just a few minutes of your time, you can teach science to thousands of curious people -- and have a permanent record of it, to boot! Below is my interview with Tobias. My questions are bolded. Interview with Tobias What motivated you to use reddit as a science outreach tool? Reddit is a fun place to goof off, but being a science nerd, I find lots of interesting science articles and images. Three years ago I found /r/AskScience, a subreddit dedicated to answering scientific questions. It is very satisfying to be a panelist there and be able to have dialog with the interested curious people asking questions about the stuff I love to think and talk about. The diverse community of panelists who answer questions, scientists from around the world, have grown to be colleagues. We share a common interest in education and outreach, and we even have lab meetings and discuss science, commiserate and share articles, skills, knowledge and excitement for science. There's also been a fair amount of collaborations struck up. It's learning from these colleagues as well as mentoring young scientists that has kept me motivated. How do you use reddit to communicate science to the public? I've been a moderator of AskScience for three years and we have built an amazing community that is reaching 3 million unique households with over 10 million page views per month! But even with all this traffic, we have a very personal relationship with our readers. People submit questions and we answer them. It's kind of simple, but what sets AskScience apart is the way we moderate the site with the help of the community. We have a team of dedicated, generous and skilled people who are cleaning up threads and deleting distracting responses. Our community helps tremendously by upvoting the best answers and downvoting or flagging problematic ones. Behind the scenes we have constant discussions about how to best serve our community and developing new ways to improve the site, reach more people, increase the signal to noise ratio. What are the pros and cons of using reddit as a science outreach tool? The pros are that it's fast easy and you can reach a wide audience. For those of us who study the science behind the things that people wonder about a lot, the possibilities are enormous. I like that you can do it whenever you want for as long as you want. Just show up and answer questions and engage in dialog. The down sides are that some panelists get burnt out answering similar questions that come up frequently. Other panelists don't get enough questions in their field or specialty. And you need to have patience and pedagogy. It can also be frustrating to see answers that sound right and get upvoted but aren't accurate. The moderators put a lot of energy into solving these problems and keeping up with our tremendous growth and traffic. What was your best experience engaging with the public on reddit? Have you had any negative experiences while engaging with the public on reddit? If so, please elaborate. Teaching people how to ask a good question is one of my favorites. Lots of questions come in that are unanswerable because they aren't scientific. Giving someone an answer to a question is satisfying but teaching them to think scientifically is even better. Learning how people think, about misconceptions and what they find interesting has proven invaluable to me and helps me communicate broadly as well as to connect with my students. It's just a pleasure to have life-long learners seek us out and learn through dialog, not only what they thought they didn't understand, but new things. We've had teachers post questions from primary school kids and those have been especially memorable. Honestly I can't think of a really negative interaction I've had with the public. Rarely people get upset because their posts get deleted, but we have an incredible amount of support from our community who appreciate the hard work we do answering questions, doing research, and cleaning up the comments. Unlike on many other sites, trolls get weeded out of AskScience very quickly. For example, we don't have creationists harassing evolutionary biologists because they get no traction on our site; they are outnumbered, downvoted, reported and anti-science comments are quickly removed. Recently sites like Popular Science have shut down their comments section because research has shown that internet trolls yelling negatively impacts learning. But we've solved that problem in a way I believe is unique on the internet. What other science outreach programs have you participated in? How would you compare science outreach on reddit to other science outreach programs you have participated in? I'm collaborating with the National Geographic Society, the Mystic Aquarium and the Mill River Conservancy to learn about snapping turtles in Connecticut. We have successfully deployed the NGS crittercam on lots of snappers over the past six years and learned things only snappers knew before now. This collaboration has educated lots of kids, allowed me to mentor many students, publish a paper and has garnered a lot of media attention. We even helped get some legislation passed to protect snappers. This work is logistically very expensive and difficult and slow. It's tremendously rewarding and successful but we haven't had grant support and can only outfit one turtle at a time with a camera that collects eight hours of footage. It takes dozens of people and dozens of hours of work for every deployment. Reddit is the opposite. Show up any time you want, stay for five minutes or five hours: no commitment, no organization, just reach out to potentially hundreds of thousands of people per day. I get different things out of these two outreach efforts and they are both very valuable to me. Do you have any advice for scientists interested in engaging with the public through reddit? It's easy. Come to /r/AskScience and check out the breadth of questions and the depth of answers and dialog. Sign up for a username and start answering questions. Once you've gotten the hang of it, apply to be a panelist and get a fancy colored tag that describes your field and specialty. If you have any questions or concerns, message the moderators and we will help you. It's a great place to have fun and do good while goofing off. Could science outreach on reddit help your career? As someone who just landed a tenure track faculty position, I found having AskScience on my CV was really valuable. Several people asked about it on interviews and this type of education, outreach and broader impacts is increasingly important to granting agencies, hiring, and tenure & promotion committees. Any parting words? If you want to challenge yourself to be an educator, learn new things and interact with the public, check out AskScience. We are growing rapidly and need the help of scientists. Big things are happening and we are heading into exciting new territories. We have just announced a new crowdsourced funding initiative open to panelists which will give our readers a new level of interaction with our panelists. Join us! Have an hour to spare? If you have some spare time and want to impart your knowledge to thousands of curious people across the world, come see if anyone at /r/AskScience is asking a question you can answer. Make sure to read their posting guidelines first, though! --- ## Discussing evolution on reddit: Interview with Bjørn Østman URL: https://www.randalolson.com/2013/11/11/discussing-evolution-on-reddit-interview-with-bjorn-ostman/ Published: 2013-11-11 Categories: outreach, reddit Tags: evolution, iama, reddit, science outreach, social media Randy Olson interviews evolutionary computational biologist Bjørn Østman about his experience giving a reddit IAmA. As a followup to my previous blog post about using reddit AMAs as a form of science outreach, I thought it'd be helpful to interview a few scientists who are already doing science outreach on reddit and discuss their experiences. If you're on the fence about whether reddit is a good place to do science outreach, give these interviews a read and see if it sounds like your cup of tea. My first interview is with Bjørn Østman, a 4th year postdoc and computational evolutionary biologist at Michigan State University. Bjørn holds a M.S. in Astrophysics from the University of Copenhagen, a M.A. in Biology from UC Santa Barbara, and a Ph.D. in Computational Biology from the Keck Graduate Institute. During his free time, Bjørn runs the Carnival of Evolution, a popular monthly review of the latest research on evolution. He also blogs about evolution and his interactions with creationists on his personal blog. @CarnyEvolution reddit IAmA Last Sunday afternoon, Bjørn ran an impromptu reddit IAmA to talk about his research and evolution in general. Following standard IAmA practice, he posted a picture verifying that he was the actual person answering the questions. [caption align="aligncenter" width="600"]Bjørn's photo verification for the reddit IAmA Photo c/o Bjørn Østman[/caption] He also provided a brief introduction to himself and his research so people knew what kinds of questions to ask. He wanted to make it clear that he was giving the IAmA to discuss evolution and careers studying evolution. Bjorn Ostman's introduction to his AMA Within 5 hours of starting, the IAmA had reached the front page of reddit, "the front page of the internet." That means that around 2.8 million people around the world logged into reddit that day and saw his IAmA (estimated from these stats). Bjorn Ostman's AMA reaches the front page After 24 hours, the IAmA was still active with users asking questions and discussing evolution. If you look at the second picture I included, users had posted over 2,000 comments focused on Bjørn, his career, and his research. A simple post and a few hours of his time had inspired thousands of users to learn more about evolution -- and some even to pursue a career studying it. Below, I interviewed Bjørn about his experience giving a reddit IAmA. My questions are bolded. Interview with Bjørn What motivated you to use reddit as a science outreach tool? This post did, actually. I really like talking about evolution, and figured that a Reddit AMA would solicit a few questions. I was overwhelmed by the interest. What are the pros and cons of using reddit as a science outreach tool? One pro is that it apparently has a lot of users and that they are particularly inquisitive. What was your best experience engaging with the public on reddit? Have you had any negative experiences while engaging with the public on reddit? I most enjoyed answering questions about my career, actually. That seemed to be helpful, I hope, to a lot of people who want to be a scientist. No negative experience with his AMA. What other science outreach programs have you participated in? How would you compare science outreach on reddit to other science outreach programs you have participated in? I blog, I do Carnival of Evolution, and I have participated in some school and high school programs through BEACON. I really like that Reddit has so many users that can be reached so easily. What was it like giving a reddit AMA as a scientist? It was intense. I answered questions for 8 hours straight, and still go back there to answer a few questions. It was a genuinely positive experience, and I'd like to do a more focused one again soon. Do you have any advice for scientists interested in engaging with the public through reddit? Yes, just do it already, but be prepared to spend a long time on it. I can only imagine how disappointing it is for people with burning questions, and then they don't get an answer. Bjørn added that creating a Frequently Asked Questions page (FAQ) as he encountered the same questions repeatedly was a major time saver. Anything else you would like to add? If I had known about Reddit AMAs before, I would have done it a long time ago. This way I can reach many more people than via blog, Twitter, Facebook, and G+ combined. Interested in running your own reddit IAmA? If a reddit IAmA sounds like something you'd like to do, make sure to read this guide on how to run one. Like any online community, reddit has its own culture and customs that are important to know about before you wade into the community. --- ## Which social media platform offers the biggest bang for the buck for science outreach to the public? URL: https://www.randalolson.com/2013/11/03/which-social-media-platform-offers-the-biggest-bang-for-the-buck-for-science-outreach-to-the-public/ Published: 2013-11-03 Categories: outreach, reddit Tags: blogging, facebook, iama, outreach, reddit, science, twitter Randy Olson outlines why he thinks reddit IAmAs are one of the better choices for science outreach on social media. Since 2005, social media has grown from a playground for angsty teenagers into a tool that pervades nearly everyone's lives. In fact, almost 3/4 of all internet-using U.S. adults now use social media in some form of another in their daily lives. It's no surprise then that scientists interested in connecting with the public have started turning to social media for science outreach. Even now, we're still very much in the "wild west" of science outreach over social media. Busy professors and scientists with little time to spare for outreach activities are likely wondering: Where can I get the biggest bang for my buck on social media? I think reddit's IAmA forum is the answer to that question. Here's why. reddit's IAmA forum If you're not familiar with reddit's IAmA forum (called /r/IAmA from now on), it's a message board where celebrities, athletes, politicians, and -- yes -- scientists go to give an open Q&A session with anyone who has a question to ask. They announce the Q&A session ahead of time, then thousands of users flock to the discussion thread to ask questions and discuss the topic or person at hand. In essence, you can have a casual conversation about your research with thousands of people at once -- all from the comfort of your office. The best part about /r/IAmA is that it's a default subreddit that every new user of reddit automatically subscribes to. Essentially, that means that reddit already has thousands of young men and women (see below) that are eager to hear what scientists have to say. Using /r/IAmA as a science outreach platform is by no means a new idea. Neil deGrasse Tyson, Bill Nye, the Mars Rover team, and many more scientists regularly show up in /r/IAmA. In fact, of the top 500 Q&A sessions in 2012, 30 of them were scientists. Science Q&A sessions can be wildly popular on /r/IAmA as well, frequently hitting the front page of reddit (dubbed "the front page of the internet"). The last time someone in my lab had a post hit the front page of reddit, the article broke every page view record ever held at the university. Last year, the Yale forum on Climate Change & the Media published an excellent blog post on the topic of reddit as a science outreach tool. There, they discuss Tony Barnston's experience running a Q&A session and provide invaluable advice for running one yourself. I highly recommend giving it a read. Why reddit over other social media platforms? reddit is by no means the most popular social media platform; Facebook wins that race. Twitter requires less time investment, and YouTube videos are more fun to watch. However, if maximum bang for your buck is what you're after, /r/IAmA offers it. Unlike the other options, /r/IAmA Q&A sessions only take a couple hours of your time and immediately reach thousands of people. There's no need to build up a web presence like you would with Twitter, blogging, and all the rest -- /r/IAmA draws the people in for you. There's no need to prepare material ahead of time -- you're just having a casual conversation about your research. And the best part? You can come back again and again as long as you have something interesting to talk about. What are you waiting for? If you're truly passionate about your science and think it's relevant to the world around you (which you should, otherwise why do it?), what are you waiting for? There's thousands of people just waiting for you to post. Organize a group together and start talking with the /r/IAmA moderators. Go on, give it a try -- but read this first. --- ## Sagemath Cloud makes collaborating with IPython Notebooks easier than ever URL: https://www.randalolson.com/2013/11/02/sagemath-cloud-makes-collaborating-with-ipython-notebooks-easier-than-ever/ Published: 2013-11-02 Categories: ipython, productivity, python Tags: cloud, ipython, latex, live collaboration, notebook, sagemath, terminal Randy Olson provides a brief overview of the Sagemath Cloud web service, focusing on the IPython Notebook live collaboration features. Earlier this year, Sagemath Cloud went into open beta so anyone can play around with it. If you're not familiar with the service, it's a free cloud-based web service that provides access to personal directories, a terminal, IPython Notebook (incl. Rmagic), a LaTeX editor, compiler, & viewer, and so much more -- all from the web, so no installation required! I see this tool being revolutionary not just for research collaboration, but for teaching as well. Imagine live coding with a class and having the code appear right on their screen as you type it. Even better, imagine not having to go through the installation process on several different operating systems every time you teach a new class. I ran across Sagemath Cloud last month and have been using it to collaborate with my mentees ever since. Since I have a bit of experience with it, I'll provide an overview of what you can do with Sagemath Cloud below. Disclaimer: I don't work for nor with the Sagemath Cloud folks, I'm just really excited about this service! All the basics Of course, Sagemath Cloud has all the basics. It has a directory structure for you to upload and organize your files and a terminal to run programs and scripts as you please. With the click of a button, you can start up an IPython Notebook and start coding away. They seem to have all of the standard scientific computing libraries installed already. I was able to replicate my Python stats tutorial without a hitch. They even have the IPython magic libraries installed, so I was able to replicate my Rmagic IPython tutorial and run R code in the IPython Notebook as well. They don't have all Python libraries installed yet, and it doesn't look like you can install your own libraries just yet. I'm sure they will install libraries by request however. It'd be especially nice if everyone could have their own Python installation and install any library as they like. Edit: Apparently it is possible to install your own Python libraries. See Aron's comment. LaTeX Sagemath Cloud lets you upload files directly from the web, so if you're writing a PLoS article, you can upload the template to Sagemath Cloud and immediately start editing it from there. You can do this with IPython Notebooks as well, so if you find a notebook on the web that you want to hack away on, just upload the file from the URL here. Even better, Sagemath Cloud has a built-in LaTeX editor, compiler, and viewer. My lab has been paying for this service at ShareLaTeX, but now it's completely free here. Live editing and collaboration One benefit of working on the web is that multiple people can work on a document at once. Sagemath Cloud makes this happen for all of the services it offers, including LaTeX. Everything has been fairly run-of-the-mill for people who are familiar with these kinds of web services so far, but the following is what really sets Sagemath Cloud above the rest for me. You can write and execute code live with others on the same IPython Notebook. My mind was blown when I saw this feature. This is a feature I've wanted for IPython Notebook ever since I started using it. The future is here, folks! Automatic backup Another benefit of working on the web is that you don't have to worry about backing up files. Sagemath Cloud takes care of all that for you and makes frequent snapshots of your working directories so you can return to previous versions of your work without a problem. Remember: Sagemath Cloud is in beta testing I'll end this with the note that Sagemath Cloud is still in beta testing. That means it won't always run perfectly, probably has some bugs, and may even crash sometimes. It's not intended for production use yet, but it will eventually have buy-in options to have priority access to the service. Make sure to report any problems you run into to William Stein. --- ## Evolving ecosystems can change more than previously thought URL: https://www.randalolson.com/2013/07/31/evolving-ecosystems-can-change-more-than-previously-thought/ Published: 2013-07-31 Categories: research Tags: Avida, digital evolution, eco-evolutionary dynamics, ecological fixed point, ecology, evolution Randy Olson discusses a recently published paper that suggests that evolution plays a large role in shaping ecosystems than previously assumed. For decades, whenever ecology researchers used computer models to study how ecosystems change over time, they often assumed that the species in any given ecosystem are more-or-less fixed. The abundances of each species may change over time -- and some species may even go extinct and be replaced by an existing species from another ecosystem -- but once a stable ecosystem is established, new species aren't going to evolve. In ecology research, such an ecosystem is said to have reached an ecological fixed point. However, recent research suggests that ecosystems may not be as fixed as previously believed. Ecological communities may look like fixed points only because new species don’t evolve on timescales that are easily observed in our lifetime. If it takes thousands or even millions of years for a new species to evolve, then a community may look fixed on the scale of hundreds of years, but be quite fluid on the scale of thousands or millions of years. The new evidence suggests that natural ecosystems exist in a dynamic steady state when studied from the point of view of longer time scales, where new species are constantly evolving and replacing the incumbent species over extremely long time periods. Of course, it's difficult to study whether natural ecosystems actually exist in a dynamic steady state because it would take thousands of years to conduct such an experiment. That's why we turned to the digital evolution platform called Avida -- where we can simulate evolving ecosystems over thousands of digital years -- to try and answer this important question. Below is a high-level summary of our findings from the BEACON class research project. If you want to get into the nitty-gritty details of the experiments, we're publishing a report on the findings in the Proceedings of the European Conference on Artificial Life in September 2013 (preprint here). Evolving digital ecosystems: dynamic steady state, not ecological fixed point To study whether ecosystems exist in a dynamic steady state, we had to simulate an evolving ecosystem for thousands of digital years. In Avida terms, we evolved the ecosystems for 500,000 updates (roughly 60,000 generations), which is more than enough time for an ecosystem to reach a stable state in Avida. At update 500,000, we took a snapshot of the ecosystem and determined the species present at update 500,000. Following that, we evolved the ecosystem for another 500,000 updates, took another snapshot, and determined the species present at update 1,000,000. What we found was surprising: the species had drastically changed after 500,000 updates of evolution! A simplified view of an evolving digital ecosystem Many species in a stable ecosystem can change after thousands of years of evolution. For example, the "green circle" species evolved into a "yellow star" species that doesn't resemble its ancestor at all. Similarly, the "blue triangle" species evolved into a "green triangle" species, which could mean that the species evolved to live off of a different food source. Finally, the "red square" species stayed the same after several thousand years of evolution; whether that was due to chance or the species was strongly selected to remain the same is a subject of future study. If evolving ecosystems do indeed exist in an ecological steady state, we should have seen the exact same species at updates 500,000 and 1,000,000. Instead, we saw ecosystems that consumed the same resources and had the same number of species, yet the individual species had changed after thousands of digital years of evolution. We interpreted this phenomenon to suggest that natural ecosystems exist in a dynamic steady state rather than a single ecological fixed point. This finding has far-reaching impacts in ecology research, especially with the recent findings that some species can evolve much faster than traditionally assumed. Population bottlenecks aren't so bad in the long run Another open question affecting ecosystem dynamics over time is whether extreme events such as catastrophic population bottlenecks have a lasting impact on the long-term evolution of ecological communities. To investigate this question, we ran the same experiment as above, except we randomly killed all but a fixed number of organisms after we took the snapshot of the ecosystem at update 500,000. We know, it's a little cruel to kill so many digital organisms en masse, but we don't have to get IACUC approval for experiments with digital organisms (yet). These experiments yielded another series of interesting and unexpected results. In the report linked above, we showed that regardless of the size of the bottleneck, the ecosystem recovers from the population bottleneck in the long term regardless of the size of the bottleneck, even if the ecosystem is reduced to a single organism. This result tells us that ecosystems are surprisingly robust to catastrophic events over long timescales. Don't worry about pandas going extinct; another panda-like species will evolve again in a million years. Photo: flickr/Chris Wieland When we compared the species at update 1,000,000 to their ancestors from before the bottleneck, we found that they were also quite different. Surprisingly, when we compared how different the species were using a species similarity measure, we found that the species in ecosystems that experienced a population bottleneck were just as different from their ancestor population as the species in ecosystems that never experienced a bottleneck. This result again tells us that population bottlenecks have a negligible impact on the long-term evolution of ecosystems, and more importantly, hints that neutral evolution plays a much larger role in shaping ecosystems over long timescales than we previously thought. This study teaches us yet another important lesson about how evolution shapes ecological communities, perhaps best captured in Bob Dylan's famous hit song: for the ecosystems, they are a-changin'... --- ## Evolution isn't over until you click stop URL: https://www.randalolson.com/2013/07/21/evolution-isnt-over-until-you-click-stop/ Published: 2013-07-21 Categories: research Tags: digital evolution, evolutionary computation, false conclusions, long term evolution, performance comparison Randy Olson discusses why it is important to make longer runs in Evolutionary Computation and Digital Evolution research. Why we need to run simulations out longer in Evolutionary Computation and Digital Evolution research I thought I had this whole evolutionary computation thing down by now. After all, I've been evolving things inside the computer for 5 years. I've evolved robots. I've evolved digital critters. Heck, I've even studied how evolution works in complex digital fitness landscapes. But nope. As always, evolution managed to surprise me. It wasn't a good surprise, either. It nearly led me to the wrong conclusion in one of my projects. This time, I learned an important lesson about evolution: Evolution isn't over until you click stop. "What does that mean?," you ask. Well, let me tell you the story of how I came to this realization. Short runs can lead to false conclusions For the past year, I've been studying collective animal behavior using digital evolutionary models. I want to get at how and why animals work together in groups. For one of these projects, I was studying the effect of predator attack mode on the evolution of swarming behavior in a selfish herd. Everything was going smoothly. I had run my evolutionary simulations out to 1,200 generations -- the normal number of generations I run my simulations out to -- and I saw some clear evolutionary trends in the data, shown below. Artificial selection experiments at 1,200 generations. "Mean Nearby Prey" is a measure of how densely the prey are swarming, or if they are swarming at all. Error bars are 95% confidence intervals. From this data, I felt confident making three conclusions: the Outside Attack treatment was strongly selecting for swarming behavior, the Random Walk Attack treatment was weakly selecting for swarming behavior, and the Random Attack treatment was not selecting for swarming behavior. Perhaps the most controversial conclusion was the third one. It completely went against people's intuition of how the selfish herd theory works. Why wouldn't random attacks select for selfish herd behavior? There has to be an advantage to swarming, even with random attacks! But the data was pretty clear: Even at generation 1,200, the prey weren't swarming at all in the random attack treatment. Even the difference in "swarminess" between generations 1 and 1,200 in the Random Attack treatment wasn't statistically significant. "Are you sure you ran the simulations out long enough?", one of my colleagues asked. I chuckled, feeling confident that 1,200 generations was more than enough to get a sense for the evolutionary trends in my model. And of course he'd ask that question. He's one of the researchers studying long-term evolutionary trends in the Long-term Experimental Evolution project. Nevertheless, I ran my experiments out for another 800 generations (for a total of 2,000 generations) to sate his curiosity. Artificial selection experiments at 2,000 generations. "Mean Nearby Prey" is a measure of how densely the prey are swarming, or if they are swarming at all. Error bars are 95% confidence intervals. That's when things got interesting. Suddenly that difference in "swarminess" between generations 1 and 2,000 in the Random Attack treatment was statistically significant. Suddenly it looked the beginnings of swarming behavior was evolving by generation 2,000 in the Random Attack treatment. My curiosity piqued, I ran the experiments out to 10,000 generations... Artificial selection experiments at 10,000 generations. "Mean Nearby Prey" is a measure of how densely the prey are swarming, or if they are swarming at all. Error bars are 95% confidence intervals. WOAH! What a difference a "few" generations can make! If you look at generations 1,000 and 2,000 for the Random Attack treatment in the graph above, what originally looked a lot like a plateau was just the beginnings of a slow ascent toward swarming behavior. In fact, the difference between generations 1 and 2,000 look nothing like a plateau in this graph; it's pretty clear that the prey were beginning to swarm at that point. Rather than not selecting for swarming behavior at all, the Random Attack treatment merely exhibited a very weak selective pressure for swarming behavior. My original conclusion was wrong because I was ending my runs too early! It was at this point that I learned... Evolution can be surprisingly slow Artificial selection experiments at 50,000 generations. "Mean Nearby Prey" is a measure of how densely the prey are swarming, or if they are swarming at all. Error bars are 95% confidence intervals. It ended up taking about 20,000 generations for the prey in the Random Attack treatment to evolve visually noticeable swarming behavior, which is roughly 20x longer than the other two treatments. While I certainly wasn't wrong in claiming that the Random Walk Attack and Outside Attack treatments select much more strongly for swarming behavior, if I hadn't run the simulations out longer, I would have wrongly concluded that Random Attacks don't select for swarming behavior. I would have missed out on something much more interesting: that random attacks do indeed select for swarming behavior, but very weakly. And that's how I learned my lesson: Evolution isn't over until you click stop. I was ending my runs too early and wasn't giving my digital critters enough time to figure out how to survive in their world. What does this mean for Evolutionary Computation research? I've read my fair share of Evolutionary Computation literature, and I'm not alone in cutting my runs short. I lost count of how many papers that report on simulations that only ran for 100 or 200 generations and claim that such and such evolutionary algorithm fails to find the optimal solution. Are 200 generations really enough to effectively explore the fitness landscape? If we're really so limited on time that we can only run 200 generations, is an evolutionary algorithm really the right algorithm to use? The problem is even worse in projects that go unpublished. How many times have researchers run an evolutionary algorithm out for a couple hundred generations, didn't seen any progress, and called it quits and moved on to the next setup? I implore all Evolutionary Computation researchers to echo the words of my colleague to themselves the next time this situation is encountered: "Are you sure you ran the simulations out long enough?" You could be right on the edge of a really interesting discovery. If your evolutionary algorithm isn't performing as expected within a couple hundred generations, don't give up. You could be right on the edge of a really interesting discovery. --- ## A short guide to using statistics in Evolutionary Computation URL: https://www.randalolson.com/2013/07/20/a-short-guide-to-using-statistics-in-evolutionary-computation/ Published: 2013-07-20 Categories: research, statistics, tutorial Tags: 95% confidence interval, benchmark testing, evolutionary computation, performance comparison, statistics Randy Olson discusses how and why Evolutionary Computation researchers should incorporate statistics into their research. A couple weeks ago, I attended the Genetic and Evolutionary Computation Conference (GECCO) for the first time. While I was perusing through the workshops and tutorials available in the first couple days of the conference, I noticed something peculiar: There was an introductory tutorial on statistics. "Why are they teaching basic stats at such a prestigious conference?," I wondered. "Shouldn't most Evolutionary Computation (EC) researchers here know basic stats and how important it is?" Once I sat through the first day of presentations, I quickly realized why that tutorial was offered: At least half of the presentations were lacking statistics! That's when I remembered that Computer Scientists aren't really taught statistics. During my undergraduate CS degree, I took an intro stats course during Freshman year that I was happy to banish from my mind after the final exam. During my graduate CS training, I learned a little bit about stats in my data mining course, but no teachers ever stressed stats as a method for inference from experimental data, nor did they talk about 95% confidence intervals and P-values. In fact, the only reason I learned about using stats was because of my interdisciplinary training in biological studies, where researchers live and die by making statistical inferences from their data. Since not everyone can attend the GECCO workshops, and assuming my CS training is representative of most CS training nowadays, I've written a short guide to explain why statistics is important in EC and how to use basic statistics in EC research. Why statistics is important in EC EC is inherently random In core Computer Science research, we rely on proofs and establishing best- and worst-case scenarios (i.e., performance boundaries) to scientifically validate and compare our work. For deterministic algorithms, that works fine. However, we have to keep in mind that EC uses an evolutionary process to guide the search, which means that the search process is inherently random to some degree. This is an important point to keep in mind because it means that two separate runs of the same evolutionary algorithm (with different random number seeds) can come up with completely different results. The random nature of EC is even more important to keep in mind if we want to compare our evolutionary algorithm to another evolutionary algorithm. If we run both algorithms only once and compare their fitness over evolutionary time, chance is going to play a huge role in which algorithm appears to perform better. For example, if our evolutionary algorithm just happens to start near the best solution in the benchmark problem we're using, and the other evolutionary algorithm starts far away in the search space from the best solution, we might conclude that our algorithm performs better than the other algorithm. Such a case is represented in the graph below. Graph showing the fitness over time for a single evolutionary run However, the opposite situation might occur, and we could falsely conclude that our evolutionary algorithm performs worse. Such is the situation in the graph below. Graph showing the fitness over time for another single evolutionary run Clearly, we can't make any comparisons between evolutionary algorithms based on individual runs. Showing the average isn't enough Some EC practitioners, privy to the above challenge, respond by running several replicates of each evolutionary algorithm (with different random number seeds) and show the average performance of each algorithm over time. This is typically what we see in EC conference presentations without statistics, where they plot the average fitness over evolutionary time. The problem here is that the graph doesn't give us a sense of the range of performance of each algorithm. Remember, each evolutionary run will give a different result, which means that multiple replicates of an algorithm will perform differently. Without a sense for the range of each algorithm's performance, we can't really compare the two algorithms. The algorithms could be all over the place, sometimes performing well and other times performing poorly, or they could always perform well. Without a sense of the range, we don't know. Graph showing the average fitness over time for 100 evolutionary runs For example, look at the graph above. It looks like our evolutionary algorithm is performing better than the current best algorithm on average over many replicates, right? Let's show the 95% confidence intervals (estimated by showing 1.96 x standard error) to give a sense of the range of performance. Graph showing the average fitness over time for 100 evolutionary runs with confidence intervals The results don't look so different after all. The confidence intervals are wide, which means that the performance of the algorithms is all over the place. There's also quite a bit of overlap between the confidence intervals, which means that there isn't really a difference in performance between the two algorithms. The reason our algorithm seemed to perform better on average was only due to chance. Statistics helps prevent false positives In short, without showing some basic statistics, we could be misleading ourselves and others who follow up on our work. In the case above, if we hadn't looked at the 95% confidence intervals, we might have mistakenly claimed that our algorithm performs better than the current best algorithm. And that, if you ask me, is worse than not publishing at all. Using statistics in Evolutionary Computation Now that we're motivated to use statistics in Evolutionary Computation, there are just a few basic concepts we need to go over. Once we have these concepts down, we're ready to incorporate a basic, yet powerful statistical analysis into our EC research. Replicates, replicates, and more replicates First and foremost, if we want to get a sense for how our evolutionary algorithm is performing, we need to run it many times with different random number seeds. A general rule of thumb is that 30 replicates should be enough to get a broad idea of the evolutionary trends, then 100 or more replicates should be performed for the setup that is going to be published. In statistics, the Law of Large Numbers tells us that the more replicates we run, the closer we get to finding the true average performance, so more replicates is always better. Standard error As I alluded to earlier, the standard error gives us an idea of the range of our evolutionary algorithm's average performance. The higher the standard error, the larger the range of the performance, and vice versa. Standard error is computed with the equation: [pmath]std error = {std deviation} / sqrt{N}[/pmath] where std deviation is a measure of how varied the fitness values are among our replicates and N is the number of replicates we ran. Every scientific computing library (R, matlab, SciPy, etc.) worth its salt has the standard error function built in, so we don't have to worry about coding it ourselves. One interesting thing to note here is that we can make the standard error incredibly small by running thousands of replicates (N = 1000s), which means that we have a highly accurate measure of the evolutionary algorithm's average performance. 95% confidence intervals Finally, from the standard error we can construct an estimated 95% confidence interval that shows us the range of our evolutionary algorithm's performance. If we were to run the algorithm multiple times again with different random number seeds, there is a 95% chance that the average performance of those replicates would fall within the confidence interval. If we want to compare the performance of two algorithms, all we have to do is check whether the algorithm's 95% confidence intervals overlap. If they overlap, then there is a high chance that they perform the same. The more the confidence intervals overlap, the more likely the algorithms are to perform the same. On the other hand, if the 95% confidence intervals don't overlap, then the algorithm with the highest average performance performs significantly better. One can imagine how this is useful for evolutionary computation research. Now we can make a statement such as, "We are confident that our evolutionary algorithm performs significantly better than the current best evolutionary algorithm at least 95% of the time," which is a scientifically powerful statement to make. Graph showing the average fitness over time for 100 evolutionary runs with confidence intervals, highlighting significant differences On top of that, because we show 95% confidence intervals, we can also think about whether our algorithm is realistically performing better than the other algorithm. For example, let's say we are comparing two evolutionary algorithms that evolve controllers for walking robots, and the graph above is a measure of how far the robots walked before falling over (in millimeters). Even if the confidence intervals don't overlap (indicating statistical significance), is walking 90mm really much better than walking 70mm? If you ask me, if the robot is only walking 90mm (~3.5 inches), neither algorithm is performing very well. Always remember to ask yourself: is the improved performance from the new algorithm realistically better? We can compute an estimated 95% confidence interval by using the following equations: [pmath]CI high = average + (1.96 * std error)[/pmath] [pmath]CI low = average - (1.96 * std error)[/pmath] where the confidence interval ranges from CI low to CI high. CI low and CI high are the minimum and maximum points that we plot for the confidence intervals in the figures above. Again, scientific computing libraries already have functions built in to calculate 95% confidence intervals (both estimated & more precise versions), so don't worry about programming it. Moving forward and changing the EC culture Computing statistics on the performance of our evolutionary algorithms is incredibly simple, yet it allows us to make powerful statements about their relative performance. We should strive to build a culture of statistical analysis into the EC research community so we can avoid false positives on performance comparisons. That said, I hope this guide has convinced you that if you write or review a paper discussing an evolutionary algorithm, you will insist that at least basic statistics such as 95% confidence intervals be included before it is published. --- ## Evolved artificial intelligence can play video games better than humans URL: https://www.randalolson.com/2013/06/09/evolved-artificial-intelligence-can-play-video-games-better-than-humans/ Published: 2013-06-09 Categories: research, review Tags: artificial intelligence, Atari, evolved AI, high score, Humies, video games Have you ever wondered if an AI could outplay you at Space Invaders? Randy Olson talks about an evolved AI that actually can. Have you ever wondered if an AI could outplay you at Space Invaders? Wonder no more. Matthew Hausknecht and his colleagues from the University of Texas at Austin report that they have evolved an AI controller that beats the highest recorded scores of several Atari games. In their report, Hausknecht et al. explain that they created the AI controller by training an artificial neural network -- a digital abstraction of how the human brain works -- to play 61 Atari games and achieve the highest score possible in all of them. They trained the artificial neural network using an evolutionary algorithm, that is, an algorithm that trains many artificial neural networks at once in a group by simulating evolution. The artificial neural networks that achieve the highest scores in that group are copied into a new group and tweaked slightly so they perform differently (for better or for worse!), then the process is repeated over and over until one of the artificial neural networks in the group can play all of the games well. Freeway, Asterix, Space Invaders, and Pitfall were just a few of the 61 Atari games that the evolved AI controller learned to play. Pictures from Hausknecht et al.'s report. After several days of virtual training, the AI controller had undergone a training montage equalling that of Rocky Balboa. While at first the best AI controller could barely even figure out how to play the games, it now stepped into the ring and took the highest scoring title for several of the Atari games. This performance is particularly impressive because it was accomplished by the same AI controller design, whereas most researchers usually have to customize their AI design for each game. Videos of the record-breaking AI controller The evolved AI scores 407,864 in Pinball, blowing the best human score of 56,851 out of the water. In Bowling, the AI controller knocks down 252 pins, a whole 13 pins more than the best human score. Even though the AI controller didn't beat the highest recorded scores for many of the games, it learned to play all 61 games well enough to achieve a respectable score. It even learned a few clever tactics, Hausknecht reports, such as "an exploitative return on Pong that the opponent can't keep up with." The evolved AI controller (green paddle on right) learned to soundly defeat the hard-coded AI in Pong by using an exploitative return that the opponent cannot keep up with. A rogue AI won't be taking over the world any time soon, but there's nothing protecting your high score in Asteroids. See more videos at Hausknecht's web page. --- ## Bigger groups make better decisions URL: https://www.randalolson.com/2013/05/18/bigger-groups-make-better-decisions/ Published: 2013-05-18 Categories: research, review Tags: collective intelligence, decision making, wisdom of the crowds Randy Olson reviews a research paper that shows us how bigger groups can make more accurate decisions. In his seminal work in 1907, Vox Populi (The Wisdom of the Crowds), Sir Francis Galton looked at the results of a weight-judging competition where participants had to guess the weight of an ox for a monetary prize. As would be expected, there were a wide range of guesses, some of which were wildly inaccurate. Much to his surprise, however, when he pooled the guesses of all 787 participants together, the crowd guessed the ox's weight nearly perfectly. From apes to ants, animals make better decisions when they work in groups than when they're own their own. What is it about working in groups that makes us perform better? In a paper published this year in the Proceedings of the Royal Society B, Dr. Max Wolf and colleagues present one way that bigger groups can make better decisions. We know that animals make better decisions in a group than by themselves, but we're still trying to find out why. Harnessing collective intelligence In their paper, Accurate decisions in an uncertain world: collective cognition increases true positives while decreasing false positives, the authors explain that groups can harness each members' unique perspective on a problem by following a specific decision-making process. First, the group members must examine the problem themselves and secretly vote "yes" or "no." (In the paper, "yes" or "no" indicates whether they saw a predator in a picture.) Next, the group members must pool their votes and vote for a second and final time, but this time they vote "yes" only if a sufficient fraction of group members also voted yes, called the quorum threshold. The result of this decision-making process is shown in the figure below. Bigger groups make better decisions: Adding group members increases true positives (accurate "yes" votes) while decreasing false positives (inaccurate "yes" votes). In this model, each group member by themselves votes "yes" correctly 60% of the time (true positive) and votes "yes" incorrectly 30% of the time (false positive). This means that, by themselves, each group member will perform pretty poorly. However, when they follow the decision-making process described above, a group with 65 members can improve their accuracy to nearly 100%! How's that for the wisdom of the crowds? Since the authors didn't include the code with their paper, I hacked together a quick version of the model and put it up on Github. Feel free to copy and reuse it as you like. Now, you may be wondering: "How the heck does that work?" How do otherwise mediocre decision makers make near-perfect decisions when they work together? Unfortunately, the authors don't go into detail on the why of this decision-making process, which no doubt means that will be the focus of future work. What the paper does explain, however, is the importance of the quorum threshold value. If you play around with the quorum threshold value in the model I linked above, you'll quickly find that this phenomenon only occurs when the quorum threshold value is in between the true positive and false positive rates. If you set the quorum threshold value outside of that range, then adding more group members can actually make the group less accurate. Models are great, but what about the real world? What the authors did next is what really made this study impressive. Knowing that their model predicts the optimal quorum threshold value to fall between a group's true positive and false positive rate, they conducted an experiment with humans to see if this prediction holds up in the real world. Sure enough, shown in the figure below, each group's quorum threshold (labeled "escape quorum") fell in between the group's true positive and false positive rate. Quorum threshold measured in humans. Figure from http://dx.doi.org/10.1098/rspb.2012.2777. The implications of this finding are profound: Not only does having an intermediate quorum threshold increase the efficacy of a group, but it appears that humans are adapted to have an intermediate quorum threshold! It will be interesting to see how this quorum threshold varies across different decision-making tasks. I also wonder if this finding will hold if we measure it in other animal species, or if this finding only applies to humans. So, what does this mean for me? All this theory and experimental findings are neat, but what does this mean for businesses, engineers, and the public at large? As it turns out, we all have to make important "yes" or "no" decisions every day. I'll highlight a couple examples from the paper, then add a couple examples of my own. Feel free to add your own examples in the comments. Doctor When a doctor examines a patient, they have to decide whether or not the patient has a disease or other ailment. Correctly diagnosing diseases (true positive) while avoiding identifying healthy patients as sick (false positive) is a vital part of a doctor's job. As such, doctors can make better decisions by independently diagnosing a patient, then pooling their diagnoses together to make a final diagnosis. Search committee Hiring a new employee can be an expensive and time-consuming process, so it's important to choose the best employee for the position. Search committees can best assess their applicants by assessing each applicant individually before discussing the applicants as a group. Engineer When analyzing the blueprints for a building, it is vital for engineers to accurately identify any potential structural faults in the building plans. Thus, design firms are better off having multiple engineers analyze the blueprints independently before holding a meeting to discuss any faults in the plans. Further, it is likely that numerous less-experienced engineers will identify more faults than a single experienced engineer. reddit I don't know if the reddit designers intended to capitalize on the "wisdom of the crowds" with the reddit model, but there's no denying that they owe their success to it. reddit delivers the most interesting content on the internet because thousands of people are voting on each link, which no doubt leads to an incredibly accurate decision of whether each link is interesting and relevant to the subreddit. The future of collective intelligence Decision making lies at the core of nearly every problem we face nowadays. A century after its discovery, there's no doubt that we still have a lot to learn from animal collective intelligence. I hope that we see more research in this field so we can better understand how and why animals work together in groups. References Wolf, M., Kurvers, R., Ward, A., Krause, S., & Krause, J. (2013). Accurate decisions in an uncertain world: collective cognition increases true positives while decreasing false positives Proceedings of the Royal Society B: Biological Sciences, 280 (1756), 20122777-20122777 DOI: 10.1098/rspb.2012.2777 --- ## How and why do animals evolve grouping behavior? URL: https://www.randalolson.com/2013/04/24/how-and-why-do-animals-evolve-grouping-behavior/ Published: 2013-04-24 Categories: research Tags: digital evolutionary model, evolution, grouping behavior, hybrid model system, krause, living in groups, ruxton Randy Olson discusses the latest research trying to understand how and why animals evolved grouping behavior. In the concluding remarks of their book Living in Groups, Jens Krause and Graeme Ruxton highlighted "understanding how and why animals evolve grouping behavior" as one of the major topics in animal grouping behavior research that would benefit from further study. Indeed, grouping behaviors are present in animals across all taxa, ranging from the microscopic bacteria to the gargantuan humpback whales. Some scientists even believe that part of the reason humans evolved such a high level of intelligence is because they lived and interacted in groups for hundreds of thousands of years. Yet, despite the omnipresence and apparent importance of grouping behaviors, we are only now beginning to understand the mechanisms underlying these behaviors. How and, perhaps more importantly, why do animals live in groups? Animals of all shapes and sizes live in groups, yet we're only now beginning to understand why. Picture credit: Bidgee Here at the BEACON Center for the Study of Evolution in Action, we specialize in looking at life from an evolutionary perspective. By taking such a perspective, we've made a number of incredible discoveries that would not have been possible if we didn't consider evolution as an important force shaping all forms of life around us. As such, I found the concluding remarks of this book particularly interesting, and worthwhile to elaborate upon. Thinking about grouping behavior in an evolutionary context In the preceding chapters of the book, Krause and Ruxton outlined many of the leading hypotheses explaining the costs and benefits of grouping behavior. Interestingly, the authors cautioned their readers that although these benefits certainly seem plausible, the hypotheses only establish "that grouping behavior would be advantageous under certain ecological conditions," but "do not address the actual selection mechanism" that could select for grouping behavior. This statement addresses one of the major pitfalls that scientists run in to when thinking about traits in an evolutionary context. Thus, the authors found it important to clarify that just because a phenotypic trait (e.g., behavior, morphological feature, etc.) is beneficial under certain ecological conditions, it does not necessarily mean that the benefit is sufficient to select for that trait over evolutionary time. Both the benefits and the costs of the trait must be considered. The benefits of grouping behavior must outweigh the costs for grouping behavior to be viable on an evolutionary scale. Picture credit: winnifredxoxo With this fact in mind, the authors asked researchers to establish experimental systems that can directly test the various hypotheses attempting to explain how grouping behavior evolves. Mind you, Living in Groups was published in 2002, so I fully expected there to be a ton of research in this area by now. Yet, much to my surprise, I only found a handful of papers broaching the subject. I've listed the papers I've found so far below. If I'm missing any papers, please email me or leave a comment and I'll update the list. List of papers directly testing evolution of grouping behavior hypotheses Author(s) Title Year Hypothesis Christopher R. Ward, Fernand Gobet, and Graham Kendall Evolving collective behavior in an artificial ecology 2001 foraging and predation Timothy C. Reluga and Steven Viscido Simulated evolution of selfish herd behavior 2005 selfish herd theory Andrew J. Wood and Graeme J. Ackland Evolving the selfish herd: emergence of distinct aggregating strategies in an individual-based model 2007 selfish herd theory Colin R. Tosh Which conditions promote negative density dependent selection on prey aggregations? 2011 dilution effect Christos C. Ioannou, Vishwesha Guttal, and Iain D. Couzin Predatory Fish Select for Coordinated Collective Motion in Virtual Prey 2012 dilution effect Randal S. Olson, David B. Knoester, and Christoph Adami Critical interplay between density-dependent predation and evolution of the selfish herd 2013 selfish herd theory Randal S. Olson et al. Predator confusion is sufficient to evolve swarming behavior 2013 predator confusion effect One interesting thing to note here is that all of these papers use some form of digital model to directly test the hypotheses. Why is that? Why can't we just use biological model systems? As it turns out, evolving behavior in biological model systems is hard. Evolution of behavior in biological model systems takes a long time. Krause and Ruxton suggested that the best biological system to study the evolution of grouping behavior in could produce four to five generations per year. At that rate, how many years would it take to evolve grouping behavior in a species that initially does not form groups? Even with an optimistic estimate of three years (from the book), that represents a significant amount of time to run an experiment that may very well be a dud. Evolution of behavior in biological model systems is difficult to control and manipulate. Anyone who has worked with live animals knows how hard it is to experimentally control every factor in the experiment. Now imagine you want to test a specific form of selection on your experimental population, such as the predator confusion effect, without any confounding effects. I feel bad for the graduate student who gets assigned that project! Evolution of behavior in biological model systems is difficult to measure. How do you quantify "groupiness" in a biological system? There have been a few impressive approaches to measuring grouping behavior in biological systems, but they always involve time-intensive video recording and analysis. Now imagine running these video analyses on an evolutionary scale, every generation, for multiple years. I just exhausted myself by merely thinking about such an endeavor. With these complications in mind, it shouldn't be so surprising that we don't see many biological model systems for studying the evolution of grouping behavior. The hypotheses explaining grouping behavior have yet to be refined enough, and refining them in biological model systems is far too expensive (both time- and resource-wise). Digital evolutionary models as experimental test beds In the past two decades, we've seen digital evolutionary models such as Avida transform into powerful experimental test beds for studying core evolutionary processes. Researchers have used these models to refine our understanding of how evolution works (e.g., how complex traits evolve), and even to make fundamentally new discoveries (e.g., survival of the flattest). As Randall Beer aptly put, The early theoretical development of a field typically involves the careful study of simpler idealized models that capture the essential conceptual features of the phenomena of interest. Such model systems have a long history in physics. For example, it was not until Galileo's consideration of such idealized situations as frictionless planes that theoretical physics in the modern sense of the word really began. The power of such an idealization is that it simultaneously makes clear a deep principle of motion (acceleration, not velocity, is proportional to force) and provides a well-defined way in which the complicating effects of friction can be understood (as an external force acting on the system). Instead of frictionless planes, we need frictionless brains. In essence, digital evolutionary models such as EOS provide the "frictionless brains" for understanding the mechanics underlying the evolution of grouping behavior. They provide a test bed to rapidly prototype and refine our hypotheses before we conduct the expensive experiments in biological systems, thereby saving countless amounts of work and money. I may be preaching to the choir here, but I feel it's important to say: It's time for grouping behavior researchers (and biologists as a whole) to abandon the antiquated notion that digital models can't tell us anything about natural processes. An evolved digital swarm from the EOS platform Hybrid digital/biological model systems Although I'm obviously critical of using biological model systems for early hypothesis testing and refinement, there has been a recent movement to merge biological and digital systems that I feel is worth mentioning. Particularly, a more recent approach coming out of Iain Couzin's lab has shown exceptional promise. In this experiment, Ioannou et al. projected virtual prey onto the side of a fish tank and had live predatory fish "feed" on the virtual prey. As the title of the paper suggests, the predator's feeding preferences selected for grouping behavior in the prey after several simulated generations. Effectively, this experiment demonstrated one method by which predation can select for the evolution of grouping behavior in prey with a real predator. How impressive is that? (Although, to make a small comment on the experiment: It only worked because grouping behavior was present in the population from the beginning, and as such does not yet explain how grouping behavior arises in the population in the first place.) Video demonstration of Ioannou et al.'s hybrid model system These hybrid model systems seem to capture the biological complexity of traditional biological systems, while still retaining many of the advantages offered by digital systems. I became so enamored by the idea of hybrid systems that I put together a proposal to build such a hybrid system myself. So, where does that leave us? Understanding the mechanisms underlying the evolution of grouping behavior is going to be a difficult yet enlightening line of research for grouping behavior researchers. Digital evolutionary models have shown promise of expediting this line of research by establishing a strong basis in theory before the experiments in biological systems proceed. Although digital evolutionary models are unlikely to make exact quantitative predictions about how grouping behavior evolves (e.g., "the predator confusion effect confuses the predator 50% of the time when 10 prey are in the group"), they will allow us to make qualitative predictions about the effects of various selection pressures (e.g., "the predator confusion effect is sufficient to evolve grouping behavior"). Once a strong basis in theory is established, we can move forward into testing our hypotheses in hybrid digital/biological model systems and eventually fully biological model systems to relax the assumptions of our models. In the meantime, however, this line of research would benefit most from concentrating on refining theory in digital evolutionary models. --- ## Is the Open Science movement prone to call-outs, and is that a good thing? URL: https://www.randalolson.com/2013/03/30/is-the-open-science-movement-prone-to-call-outs-and-is-that-a-good-thing/ Published: 2013-03-30 Categories: open science, philosophy Tags: arsenic life, call-out, career, damage, donglegate, open science Randy Olson ponders if the Open Science movement is prone to call-outs, and whether that can be damaging to scientist's careers. I had a discussion with a couple scientists this week who are critical of the Open Science movement. Among the usual counter-arguments, one argument stuck out to me as particularly thought-provoking: (paraphrased) Open Science is prone to call-outs, and that can be unnecessarily damaging to scientist's careers. In light of the recent #DongleGate travesty, we've seen how damaging public call-outs can be to the careers of everyone involved. No one wants to be affiliated with a person who is actively protested against on the Internet. It's bad for publicity, and therefore it's bad for business. The Tweet that was heard around the Internet (Image thanks to @adriarichards/Twitter) Call-outs in science Not too long ago, the scientific community had its own DongleGate with the #ArsenicLife controversy (overviewed by Jonathan Eisen here, which summarizes summaries by other people who summarized the issue -- yeah, it was a big thing). In the beginning, Wolfe-Simon et al. published a controversial paper in Science. Some scientists disagreed with the findings, and pursued an ideal Open Science approach to science and posted their critiques on arXiv and via social media outlets. And as Jonathan Eisen puts it, "the critiques snowballed and snowballed" and "some of the critiques got way too personal." In the end, scientific truth won out and Wolfe-Simon et al.'s work was proven flawed. But at what cost? In traditional scientific venues, controversy is a way of life. For about every stance you can take in science, you can be guaranteed that there will be someone who disagrees with you. Thus, whenever controversy arises, it's just another day in science: we look at the issue, deliberate the data presented, decide whether the data is convincing, and identify future experiments if we find the data unconvincing. And here's the most important step after all that is done: we move on with our lives. Controversy can easily snowball when news and social media gets involved (Image thanks to Kamyar Adl) News and social media venues don't work that way. Both of these venues actively seek controversy and thrive off of it. The more people that get involved, the more successful the blog posts and news articles are, and the more blog posts and news articles are created to feed off of the controversy. As we saw with the #ArsenicLife controversy, news and social media bloated the #ArsenicLife debate to national proportions and kept the controversy on everyone's minds every day until the issue was finally settled. What costs would we pay for Open Science? A colleague of mine related a story of how Rosie Redfield, one of the fiercest critics of the arsenic life work, gave a talk at the Evolution Ottawa 2012 Conference. Rosie basically outlined how the #ArsenicLife controversy proceeded, and further dragged Felisa's name in the mud in front of the massive audience who eagerly listened to Rosie's story. All to the uproarious applause of the audience, as Rosie was touted as the ideal Open Scientist. And understandably so: she had done great work leading the charge in the #ArsenicLife controversy. I met Felisa at the Gordon Research Conference on the Origin of Life in 2012, and I don't think I've ever met someone who was so visibly stressed out from their work. If you've followed up on what happened after the #ArsenicLife controversy, you'll know that Felisa was "effectively evicted" from her position at the USGS laboratory and forced to find a new home for her research, all because of the fuss raised in the name of Open Science. She's still an early career scientist, yet her career is marred by the #ArsenicLife controversy for the rest of her life. The desire to be open in science can potentially harm people (pictured: Felisa Wolfe-Simon) (Image thanks to NASA) With that in mind, ask yourself: Is expedient scientific truth worth damaging people's careers? Disrupting their livelihood? If the #ArsenicLife controversy is any indication, news and social media venues are not ready to handle scientific controversy properly. If we open scientific debate to everyone, the conversation won't just be limited to scientists. Do we really want a repeat of #ArsenicLife? Let's think things through before we get caught up in idealistic visions It's easy to get caught up in the Open Science notion of "being a good scientist." It makes sense. It sounds right. It feels right, on an instinctual level. All I'm asking is for us to seriously consider both sides of the coin--both the professional and the personal implications--before we change the way we do science. --- ## Could the reddit model replace the current scientific publication system? URL: https://www.randalolson.com/2013/03/23/could-the-reddit-model-replace-the-current-scientific-publication-system/ Published: 2013-03-23 Categories: open science, philosophy, reddit Tags: open access, open science, reddit, scientific publishing Randy Olson discusses how we could replace the current scientific publishing model with a model similar to reddit. Just today, Subhajit Ganguly put up a blog post that piqued my interest: "Imminent Changes In The Publication Process In Sciences." In this blog post, Subhajit repeats the battle cry we so often hear from the Open Science movement: traditional for-profit journals have outgrown their usefulness, they benefit only a few individuals, and it's time to abandon traditional journals and move to a completely Open Science model, such as posting all of your scientific output on arXiv or figshare for everyone to see and critique. The benefits of the Open Science model There are clear benefits to openly posting all of your scientific output instead of relying on for-profit journals: 1) We open the possibility of receiving feedback and criticism from everyone in the scientific community, and not just a select few who were chosen to review the article. Presumably this will lead to a more fair treatment of everyone's scientific output, since the rating of a manuscript will be based on the opinion of hundreds (or even thousands) of people with a broad range of viewpoints and backgrounds. 2) We cut out the middleman (the for-profit publisher). We won't have to pay someone to disseminate our manuscripts any more. Most scientists find and retrieve other scientist's manuscripts online nowadays. Why are we paying someone thousands of dollars to host our manuscript(s) on a web server when we can host them on figshare/arXiv for free? Then we can freely promote our article through standard means: conferences, press releases, blog posts, Twitter, etc. 3) We increase the visibility of our scientific output. There has been a big uproar in the Open Science community claiming that "hiding your scientific output behind a paywall is immoral." Posting your manuscripts on figshare/arXiv immediately alleviates this concern. Additionally, anyone can access your manuscript(s) at any point in time since they're just a click away on figshare/arXiv, which directly translates into increased potential visibility of your scientific output. I'm sure there are many more benefits, but those are the major ones that stick out to me. Let's be real My biggest criticism of moving to the model that Subhajit et al. argue for is about benefit #1. It assumes that we have hundreds (or even thousands) of scientists chomping at the bit to thoughtfully critique everyone's scientific output. Realistically, journals already have enough trouble finding just three scientists to thoughtfully review each manuscript, and that's with the added benefit that the reviewers can claim they performed a service to the journal and scientific community on their CV. If we take that benefit away, will we really average three or more reviewers per manuscript? Beyond the above, is it really practical to rely on the standard means of dissemination to get your article noticed? Publishing a manuscript in high-profile journals is a reliable way to ensure that your manuscript is seen by a good portion of the scientific community. If we take away the journals, what do we have left to make your manuscript stand out among the rest? We have some popular blogs, such as Haldane's Sieve, that highlight quality manuscripts and give them the attention they deserve. But what if the people running those blogs don't think your manuscript is good enough to highlight in their blog? What about the thousands of manuscripts that would fall into that category? We run into the exact same problem as we have with for-profit journals: when only a few people are in charge of deciding which manuscripts are "highlight-worthy," not all manuscripts are going to be given a fair rating. Could the reddit model replace the current scientific publication system? For those of you who are unfamiliar with reddit, reddit is one of the most-visited social news web sites in the world. reddit has been tremendously influential on the web as the self-proclaimed "front page of the Internet." The reddit model works as follows. Users who sign up on reddit are given four abilities: submit links comment on links and other comments upvote links and comments they think are constructive, and downvote links and comments they think are unconstructive. Links and comments that receive a higher score (score = upvotes minus downvotes) are ranked higher in the queue when people view the list of links on reddit. Typically high-scoring links are only ranked high for 24 hours after they are posted, after which time they rapidly decline in the queue. Top 5 posts on /r/science on March 23, 2013 Shown above, reddit already has an active scientific community in /r/science. There, users submit links to scientific articles, upvote or downvote them depending on their quality and relevance, and discuss their methods and implications in the comments. Seeing this community got me thinking: could the entire scientific community share research with each other this way? With this model, there's no more need for formal peer review. Instead, peer review is done by upvoting or downvoting the link and discussing the manuscript in the comments. The constructive comments would be upvoted to the top, while the unconstructive comments fall to the bottom. The entire process is transparent to anyone visiting the web site, and the discussion is archived if anyone ever wants to find that comment that Dr. So-and-so made about a manuscript four years ago. A few subreddits that you can subscribe to on reddit There would be separate "subreddits," or sub-forums for each field of research that you could optionally subscribe to. If you're only interested in sequence alignment algorithms, there could be a subreddit dedicated specifically to manuscripts detailing the latest sequence alignment algorithms. Similarly, if you have broader interests such as Evolutionary Biology in general, there could be a subreddit dedicated to discussing the major breakthrough manuscripts in Evolutionary Biology. The research community would be able to create and abandon these subreddits as they wish, allowing the system to reflect the interests of the research community as a whole. The citation system would work as usual, except the journal would be replaced by the preprint manuscript hosting service you use. Instead of reporting on your CV that you published two papers in such-and-such high-profile journal, you would report the Sciddit ("Scientific reddit") score your manuscript received. The reddit model seems like a much more scalable model for the widespread dissemination, peer review, and discussion of scientific output online. The only thing left is to give it a try. Will someone step up to the challenge? Update (03/25/2013): It looks like someone already created a reddit interface for arXiv. It's a little buggy, but this is a great first step toward making this idea a reality! http://arxaliv.org/ --- ## A data-driven guide to creating successful Reddit posts URL: https://www.randalolson.com/2013/03/15/a-data-driven-guide-to-creating-successful-reddit-posts/ Published: 2013-03-15 Categories: analysis, data visualization, reddit Tags: analysis, front page, reddit, top posts, upvotes Randy Olson provides a data-driven guide to making your Reddit posts visible to the broadest audience possible. Edit: This article has been updated to 2015. Please see the latest version of the article here. Today I'm going to tackle the question that's on all Redditor's minds: How do I get a highly-upvoted post on Reddit? I won't bother covering the basics of making a successful post on Reddit because there are at least a dozen other articles out there that already covering that. Instead, I'm going to walk you through my analysis of over 850,000 top posts from the past year on Reddit from 4,200 of the most active subreddits. I bolded the big takeaway messages if you're not feeling like a long read. Give these tips a try for a week or two and report back how well they worked for you in the comments. Disclaimer: I am only making statements about probability in this post. Following these guidelines will by no means 100% guarantee that you will get a top post. Rather, following these guidelines will maximize your chance of getting a top post. When to post It turns out one of the biggest factors affecting the success of your post is the day and time you submit it. In the top graph (below), the shaded area gives an indication of the number of upvotes I am confident a successful post will receive given that it is posted at a given day and time. Similarly, the bottom graph shows the total number of posts that will receive more than 100 upvotes given the day and time they are posted. Average number of upvotes and submissions by day of the week and hour of the day Here's a handy time zone conversion chart so you can convert the UTC time to your local time: http://www.onlineconversion.com/timezone.php Every day around 12 PM to 1 PM UTC (7 AM to 8 AM EST), there is a highly significant spike in both the number of successful posts and the total number of upvotes those successful posts receive. Undoubtedly, this trend is due to office workers in the U.S. coming into work and catching up on Reddit before they start their workday. The key to success here seems to be to (1) post early in the morning before the deluge of new posts comes in and (2) take advantage of your post's head start and get upvoted out of the /r/new queue before everyone else. If you've posted around 12 to 1 PM UTC and your post escapes the /r/new queue, then your post is in prime position to be viewed and upvoted by the U.S. workforce coming in to work in the next few hours. Interestingly, more posts are successful during the weekdays, but the successful posts on the weekends generally receive far more upvotes. What this means for you is that you stand more of a chance of having a successful post on Monday than you do on Saturday, but if your post is successful, it won't receive as many upvotes on Monday than it would on Saturday. Here's what I find amusing in this graph: the number of successful posts peaks on Mondays and then gradually declines over the week, showing that Redditors likely spend more time on Reddit early on in the week when they're suffering from a bad case of the Mondays. (And Tuesdays... and maybe even Wednesdays.) What to post Great, so now you know when to submit your post, but... what kind of content should you post? First off, let me get this one out of the way: Reddit is married to imgur.com as an image hosting service. Fraction of top posts that are imgur.com links by subreddit If you're submitting an image post, upload your image to imgur.com and submit that link. Nearly 60% of the top posts from the past year were some form of image or GIF hosted on imgur.com. With that out of the way, let's move on to the next-most-popular content format. Fraction of top posts by domain If it's not images that you're looking on the top posts page, then it's most likely a YouTube video. Reddit is dominated by image and video content nowadays. In fact, if you look at just the data presented here, at least 2/3 of the top content on Reddit is an image or video. As such, if you have some sort of message you want to share with Reddit, you're best off trying to communicate it through an image or video. quickmeme.com appears to be Reddit's favorite meme generating website, whereas blogspot.com seems to be the most popular blogging service. What I find surprising is that Facebook and Twitter links appear to be shared equally, even though many subreddits have implemented rules against sharing direct Facebook links. Meanwhile, the Wikimedia/pedia services are Reddit's go-to web sites for free educational content, showing just how much Redditors rely on wikis for their information. Lastly, Reddit's top news posts most often come from BBC, The Guardian, and The Huffington Post. When it comes to news sources on Reddit, it looks like the UK has the U.S. trumped! Average number of upvotes by domain If it's upvotes you're after, then I have a different story for you. Whereas imgur.com links are by far the most likely to rise to the top, successful /r/AskReddit self-text posts and meme posts are significantly more likely to receive more upvotes. It's undeniable that Reddit loves sharing stories and jokes via /r/AskReddit, and the patrons of /r/AdviceAnimals freely hand out upvotes to any meme that makes them so much as chuckle. Keep this in mind if you're more concerned about upvotes than getting your message out there. Note: quickmeme.com has recently been banned from reddit entirely for vote manipulation, and imgur.com was quick to make a meme generator service to fill the void. The final thing I'll note here is that imgur.com ranks 1st in terms of likeliness to reach the top page, and 4th in terms of upvotes. If you're following along, that means your best bet of reaching the top page and reaping sweet, abundant karma is to submit an imgur.com link early in the morning on the weekend. Where to post Alright, so now you know what and when to submit your post to Reddit. Where should you post it? Fraction of top posts by subreddit Surprisingly, the default image subreddits don't completely dominate the top posts. As this graph shows, the default image subreddits comprise only about 7-8% of the top-ranked posts from the past year. Let this be a heartening discovery: you don't have to post in a default image subreddit to have a successful post on Reddit. The rest of the subreddits not listed here all accounted for only ~0.25% of Reddit's top posts (each), so if you don't pick one from this list, your chances of having a successful post are more-or-less the same (i.e., low). Average number of upvotes vs. fraction of top posts by subreddit However, again if it's karma you're after, I have a very important addendum: the fraction of top posts that come from a subreddit is highly predictive of the number of upvotes your post will receive if it's successful. Successful posts in /r/funny, /r/AdviceAnimals, and /r/pics by far receive the most upvotes on average. So keep that in mind: the higher you go on the "Fraction of top posts" chart (i.e., smaller fractions), the fewer upvotes your post will potentially receive. The good news is that smaller subreddits have fewer people posting to them, so even though you receive fewer upvotes, your post is far more likely to show up on the front page of anyone subscribed to that subreddit. (For you stats nerds out there: I also fit this model with all 4,200 subreddits and the relationship is still significantly predictive, even if I remove the extreme outliers on the right of the graph.) What title to use Now for the last of the five W's: Why? (If you didn't catch it, the first W was "Who?" The answer to that one is you!) Why should someone bother clicking on your post, read it, and upvote it? That's the purpose of your post's title. Below is a word cloud of the most-used words in the top posts on Reddit. The larger the word is, the more often it was used. Most-used words in submission titles I'm going to take 40 or so of the most-used words, break them down into categories, then give you some example titles that the words were used in. Please take these example titles as just that: examples. Use them as inspiration to create your own post titles. x-post "x-post" is the most-used word in the top posts from last year, and if you're familiar with Reddit, this should not be a surprise. Cross-posting is huge on Reddit, especially when there are many subreddits with similar interests. Generally, cross-posted links do better because they're already well-liked in another subreddit before they were cross-posted, and thus were vetted for the poster beforehand. A word of warning: be careful with cross-posting to excess. Reddit can get pissed off pretty quickly if you share a link too much. Evil Elementary School Girl (x-post from r/pettyrevenge) Dashcam footage of the 2011 Earthquake and Tsunami (x-post r/roadcam) Props to the photographer [x-post from r/owls] Speaking of time... Some mention of time is commonplace in top Reddit posts. These words seem to be used to make the post relevant to current or past events, and thus make for a more interesting post title. By the time they reached the front of the line the baby had fallen asleep. Thankfully Santa played along! My Mom helping me through a hard level in Super Mario Land on the day Nintendo Gameboy was released I've been playing the same game of Civilization II for almost 10 years. This is the result. OSU marching band does a fantastic tribute to classic video games at halftime of last night's game Neil deGrasse Tyson is serving jury duty this week Cable Industry Finally Admits That Data Caps Have Nothing To Do With Congestion: 'The reality is that data caps are all about increasing revenue for broadband providers -- in a market that is already quite profitable.' So I saw this guy the other day... Mentions of other, non-specific people seem to be extremely common among the top posts. Reddit apparently likes to gossip and share stories about other people. My friend calls him "Mr Ridiculously Photogenic Guy" My friend is a college professor Male bus driver uppercuts a girl after she attacks him on his bus. I have been punished for upgrading my PC. This is why people pirate games. Man has a conversation with his 12 year old self [3:47] A group of Delhi women have decided the best way to make sure women are safe is to create a special cab service just for them. Women drive the cabs and only pick up women passengers. Puppies and kitties everywhere Reddit isn't just all cat lovers any more: cats and dogs are mentioned equally in the top posts from the past year. It's easy to relate to other people through likable animals like cats and dogs, especially when you're sharing a cute picture of them. So I walked into the kitchen at 5:30am and saw this in the sink... This is not my cat.. Dog teaches puppy to use the stairs. Gamers unite Big surprise, Redditors like to talk about games. Video games, specifically. Why don't all games do this?! (Max Payne 3) This is why it takes me so long to finish a game. Talk about what you're posting As I mentioned before, images and videos is the primary content on Reddit. It turns out it helps if you talk about the image or video you're posting. My local humane society posts pictures of new adoptions. This one makes me so happy. Words can not describe how much I love this pic of Obama and Clinton When I see a funny post with only a few upvotes So I have combined these two gifs. The story is now complete. I meticulously tabbed this video of 100 guitar riffs. Press Play and enjoy. Photo I took of the Chicago skyline from a beach in Indiana. Have a broader perspective People come on Reddit to escape their daily life and learn something about the world. If you can fill that need for them, they will upvote you for it. World's fastest archer - Reinventing the fastest forgotten archery. North Korean defector draws life in a concentration camp. I've taken the liberty to translate Emotions make people like your post more If you're really passionate about what you're posting, let it shine through! Redditors love it when a post title shows emotion. A young Steve Irwin gets bitten by a snake on television. I love the classy and calm manner in which he responds. As someone who is fairly new this is how I feel about my budding Reddit addiction... I don't care how many times people have posted this. It's still my favorite thing on the Internet. Came out of the restaurant and saw this. It made me happy. This guy is a reporter on Fox 2 here in Detroit. His name is Charlie LeDuff. He is fucking awesome. A diminutive preoccupation Reddit loves to talk about little things. I guess it's because little things are cute. I live by the beach and this little guy just popped by for a visit Walked in on my little cousins... Talk about Reddit! Reddit is big on self-referential jokes. If you can get in on those, you're golden! Screenshot of reddit from the year 3012 When I post something that Reddit doesn't seem to like Location is key If it's relevant, mention a location related to your post. It helps give context to what Redditors are looking at. Amusingly, "work," "school," and "home" are the most-used words related to location. Oktoberfest: Redditor from Munich here! This is my walk to work in the morning... Guy at my school attempting to release a mouse into the wild. Walking home last night in Brighton and turned the corner to see this tribute to MCA by graffiti artist 'Aroe'. It's unbelievable. 12ft by SIXTY FOUR FT! That's all, folks! Don't wait for chance to favor you; make chance favor you. --- ## Retracing the evolution of Reddit through post data URL: https://www.randalolson.com/2013/03/12/retracing-the-evolution-of-reddit-through-post-data/ Published: 2013-03-12 Categories: analysis, data visualization, reddit Tags: analysis, community, reddit, submission distribution, subreddit Randy Olson explores how the Reddit community has changed over time since its inception in 2005 by analyzing posts that were submitted to the various subreddits. Have you ever wondered how the Reddit community has changed since it started in 2005? I've only been active on Reddit since late 2011, so I've always been curious how Reddit looked before I came along. In this post, I try to retrace the evolution of the Reddit community from its inception in 2005 up to November 2012 by analyzing the posts that users submitted to Reddit over time. The current state of Reddit What does the Reddit community look like right now? Where do people submit their posts to? The graph below answers those questions at least in part by showing the number of posts that were submitted to each of the most active subreddits in the first week of November 2012. Distribution of posts submitted to Reddit on the first week of November 2012 If you're familiar with Reddit, the top 5 subreddits are to be expected: /r/funny, /r/AskReddit, /r/AdviceAnimals, /r/pics, and /r/aww. Four of them are focused on sharing pictures, and /r/AskReddit is used to ask questions of and share stories with other Redditors. Together, these five subreddits account for the majority of the content on the Reddit default front page nowadays. The next 6 subreddits are a little more interesting. /r/trees (a subreddit dedicated to discussing marijuana), /r/politics, /r/Music, and /r/gaming capture the main interests of most Redditors, who are typically 20-something American male college students. Finally, /r/videos and /r/WTF are two more media-based subreddits. From this brief survey, it becomes abundantly clear that the primary content of Reddit nowadays is pictures and videos. This trend makes sense, too: pictures are easy content to produce and take only a few seconds to look at, enjoy, and upvote. For better or for worse, Reddit's karma system favors pictures and short videos, and I expect there to be an ever-increasing trend of more pictures and fewer text-based posts. I should note here that there are thousands of active Reddit communities nowadays, and I have made some broad generalizations here based off of the most popular subreddits. Put away the torches and pitch forks, please! How Reddit has changed over time Now that we know what the Reddit community looks like today, how did it look in the past? The graph below shows how 24 of the most active subreddits have changed over time. I ordered the subreddits by the time that they first appeared on Reddit. I recommend zooming in so you can see it better. (I should note that I purposely excluded /r/reddit.com from this graph because it dominates the entire graph until about 2008, then screws things up again when it got closed down in late 2011.) Subreddit growth over time measured by fraction of posts submitted to each subreddit. /r/reddit.com was excluded for visualization purposes, which is why it looks like everything was in /r/NSFW at first. To clarify: reddit did not start as a 100% NSFW web site! The biggest thing that you may notice is that there were very few subreddits from 2006-2008. In fact, there was only one subreddit before 2006 (/r/reddit.com). The majority of the content in 2006-2008 was focused on more techie-friendly subjects: programming, science, politics, entertainment, and gaming. Major subreddits dedicated to solely picture and video content started becoming popular in mid-2008, and even then their posts only comprised less than 1/4 of Reddit's content. It wasn't until 2011 did the picture-related subreddits really start taking over, and Reddit never looked back after that. This graph covers so many changes in the Reddit community that it can't explain what happened by itself. In the following sections, I will take a closer look at how the Reddit community evolved on a year-by-year basis. 2005 - The Reddean Eon Reddit didn't show any signs of life until Thursday, June 23, 2005, when Reddit co-founder kn0thing submitted the first post on record to Reddit discussing The Downing Street Memo. How appropriate that the first public post to reddit was politically-oriented! Update (03/26/2013): reddit team member jedberg pointed out that there was one post by kn0thing just before the one ChickenLittle posted. Where'd that post come from!? 2005 - The Reddean Eon There wasn't really anything interesting happening during this time period, but I had to include it for posterity. At this point, Reddit only had a single subreddit: /r/reddit.com. This period marks the early development of Reddit by Steve Huffman and Alexis Ohanian, when they added core features like the "hot" page ranking (which is now default for the front page) and persistent upvotes. 2006 - The Archreddean Eon 2006 is finally when things get interesting! Presumably after one too many complaints about getting caught looking at pictures of breasts "on accident" on Reddit, they added the very first subreddit specifically for NSFW content: /r/NSFW. As you can see in the graph, people immediately started using it, showing how eager Redditors are to share NSFW content. It's amusing to read their curt blog post about making the new subreddit, where they state: We'll be playing around with more of these sub-sections to reddit in the future. Should be fun. I guess they didn't know that they had stumbled into an idea that would later make Reddit what it is today. 2006 - The Archeddean Eon You'll notice in February, the number of subreddits suddenly spikes. Sure enough, following the Reddit blog, they started adding all kinds of subreddits during February: Olympics, feature request, and subreddit request subreddits; language-specific subreddits; and yet more language-specific subreddits. A few more topical subreddits showed up later as well, which are vaguely mentioned in the blogs. For the most part, though, Reddit's content came from /r/reddit.com, which acted as a catch-all for everything else. Notable subreddits born this year /r/programming, /r/science Notable subreddits died this year Most of the language subreddits died before they even had a chance. Survival of the fittest! 2007 - The Protereddoic Eon 2007 marks the extinction of all but two of the remaining language-specific subreddits that hadn't already died. The main techie-friendly subreddits (/r/programming and /r/science) thrived, while /r/lipstick.com (a celebrity gossip subreddit) seems to have been supplanted by a more general /r/entertainment subreddit. 2007 - The Protereddoic Eon The biggest change in subreddits that we can see here is the explosive emergence of /r/politics. It seems like Redditors in 2007 were dying to have a politically oriented subreddit, seeing how the subreddit grew to take up 20% of the total submissions to Reddit in a matter of a few months. Sure enough, looking at the blogs again, we can confirm just that: /r/politics was an addition by popular demand. Redditors of 2007 were so politically-oriented that they had /r/obama created over a year before the election! Notable subreddits born this year /r/politics, /r/entertainment, /r/business, /r/gaming, /r/gadgets, /r/sports Notable subreddits died this year The rest of the language subreddits except for the Japanese and Italian subreddits. They really liked Reddit! The rest of the subreddits dedicated to outside blogs: /r/lipstick.com, /r/joel. 2008 - Reddit's Cambrian Explosion If you've been following the graphs, January of each year always seems to bring a whirlwind of change to Reddit. 2008 was no exception. 2008 - Reddit's Cambrian Explosion This year, we see an explosion in the diversity of subreddits. What the heck happened? As it turns out, Reddit added a feature for users to create their own subreddit. Almost immediately, people created /r/pics, /r/funny, /r/WTF, and many other subreddits dedicated to pictures. In just the matter of a few months, 1/3 of Reddit's content was sorted from /r/reddit.com into subreddits dedicated to specific topics. Thanks to the new "create your own subreddit" feature, Redditors were able to sort the majority of their own content out into the proper forums. The Reddit system works! Cool! It's also pretty cool to see the spike in posts to /r/politics and /r/obama near election time in November, showing just how active Reddit was in the Presidential election. It seems even from the early years, Reddit was a politically-oriented web site. I was interested in just how many subreddits were created in this time period, so I graphed "subreddit diversity" over time. Subreddit diversity over time Panel A shows the very gradual increase in subreddits from 2006-2008 as Reddit's managers slowly added subreddits by request. Panel B shows the huge blip in the number of subreddits during the "create your own subreddit" beta period, then the crash when they ended the beta period, then finally the never-ending diversification of subreddits as Redditors were officially allowed to create their own subreddit. As of November 2012, there were as many as 14,000 different subreddits active at once, and I'm sure it's kept growing since then! Notable subreddits born this year /r/AskReddit was born on January 25th, 2008 when some Redditors noticed that "Ask Reddit" was becoming a popular posting trend. Sadly, the only questions asked that day were questions about how the subreddit would work. /r/gonewild was born on December 19th, 2008 as an awkward conversation between a bunch of guys talking about showing each other their penises. Oh, internets! /r/pics, /r/funny, /r/environment, /r/worldnews, /r/Libertarian, /r/technology, /r/WTF, /r/offbeat, /r/space, /r/reportthespammers, /r/Atheism, /r/Christianity, and many many more. Sorry if I didn't list your favorite subreddit that was born during this time -- there's so many! Notable subreddits died this year Surprisingly, none! Though I'm sure many subreddits tried to start up and died during this period. 2009 - The Great /r/reddit.com Spike By the end of 2008, over 2/3 of Reddit's content had been sorted into specific subreddits, which is quite impressive considering 80% of Reddit's content was being posted to the catch-all /r/reddit.com just a year and half earlier. 2009 seemed to be much of the same of 2008, with yet more content filtering into subject-specific subreddits and /r/reddit.com becoming increasingly obsolete. 2009 - The Great /r/reddit.com Spike Then in June, something weird happened: a huge spike in /r/reddit.com posts! I've looked all over the blog and scoured the Internet and can't find a reasonable explanation for this spike. Do any Redditors from 2009 know why? What's not shown in this graph but shows up in the data: even up through 2009, there were very few pornography-centric subreddits. Either /r/NSFW and /r/gonewild handled most of Reddit's needs at the time, or Reddit just wasn't the place for pornography in 2009. Notable subreddits born this year /r/IAmA was born on May 28, 2009 with a series of IAmA posts by people from all walks of life. /r/todayilearned was born on January 2nd, 2009 with a post about how carrots were originally purple, which I am fairly positive has been reposted at least once. /r/trees was born on October 15th, 2009 with a series of fantastically on-topic posts. /r/fffffffuuuuuuuuuuuu, /r/circlejerk, /r/DoesAnybodyElse, /r/listentothis, /r/tf2, /r/TwoXChromosomes, /r/SuicideWatch Notable subreddits died this year /r/gossip, /r/pornography (a short-lived effort to share pornography on Reddit) The final language-specific subreddits originally created in 2006, /r/it (Italy) and /r/ja (Japan), died off in favor of country-specific subreddits created by the users. 2010 - The Year Gamers and Porn Invaded If it wasn't obvious in 2009, you can't miss the fact that this is the time period Reddit became less techie-focused and more like it is today. /r/politics was supplanted by /r/AskReddit and the picture-related subreddits, while all the stoners crowded to /r/trees to discuss "fucking trees, guys." Interestingly, there appears to be a decline in all of the techie-friendly subreddits at the end of 2010, when the Digg exodus was in full force. /r/gonewild also seems to have started growing in popularity during the Digg exodus (not shown here -- it was still very small), so maybe Redditors have something to thank Digg for! 2010 - The Year Gamers and Porn Invaded Meanwhile, /r/gaming and gaming-related subreddits continued to grow in popularity throughout 2010. What's really neat is this is the first year you can see subreddits exploding in size over a short period of time because of the release of a game: /r/starcraft for StarCraft II in June 2010 (hype for the release of the game the next month) and /r/Minecraft in July 2010 (when it started getting international media attention). Apparently spammers became a huge problem on Reddit in 2010, since reports on /r/reportthespammers comprised at least 10% of Reddit's submissions on a weekly basis. Thanks, Obama. Notable subreddits born this year /r/ColbertRally had a short-lived spike in the months immediately before and after the rally happened. It's really neat how you can detect the big political events that Reddit participates in by looking at these spikes! /r/leagueoflegends, /r/wow, /r/buildapc, and many more gaming-related subreddits. Notable subreddits died this year /r/RapidshareList died as quickly as it came to life, peaking out at over 5,000 submissions in the week before it died (as many as the top subreddits at the time). Did someone get busted? /r/Marijuana went down in flames as /r/trees took over. 2011 - The Year /r/reddit.com Died Ah, 2011. The year Reddit lost an old friend. After seeing /r/reddit.com still host 20% of Reddit's content after two years of trying to get rid of it, I'm sure the Reddit administrators got sick of waiting for /r/reddit.com to be replaced by topic-specific subreddits and finally decided to pull the plug. While it was sad to see /r/reddit.com go -- it was a great catch-all subreddit, after all -- we have been afforded a rare opportunity to see how removing a massive subreddit can affect an online community. Now, let's take a look... (I should note that I moved /r/reddit.com to the top of the graph here so its removal at the end doesn't cause a wave in all the other subreddit's areas.) 2011 - The Year the /r/reddit.com Died For the most part, Reddit's transition to picture-focused content continued just as it did in 2010. Memes starting sprouting up all over Reddit: rage comics (/r/f7u12), First World Problems, Advice Animals... many of our favorite memes first appeared on Reddit in 2011. Fewer people talked about politics, world news, etc. and instead preferred the easy laughs and upvotes that memes offered. Using stattit.com's subreddit time machine, we can see how dramatically the front page shifted during this time period: front page from January 1st, 2010 (only 27/100 of the top posts were images, and only a few of those are memes) front page from January 1st, 2011 (60/100 of the top posts are images, and many more are memes) If you want to see something really interesting, check the front page on January 1st, 2009 (11/100 top posts are images, no memes), or even the front page on January 1st, 2008 (1/100 top posts are images, no memes). In late October, /r/reddit.com was shut down for good and Reddit's community shifted dramatically because of it. It looks like /r/pics experienced a temporary bloat in submissions (likely pictures coming from /r/reddit.com), then suddenly the number of submissions to /r/pics shrunk. Shortly thereafter, the images from /r/pics moved to the other image subreddits (/r/funny, /r/f7u12, /r/AdviceAnimals, /r/WTf, and /r/aww) for some reason. Apparently, the /r/pics mods decided they didn't like the sudden influx of meme pictures to their subreddit, and consequently cracked down on submissions and pushed them off to more specialized subreddits. /r/funny and /r/AdviceAnimals seemed happy to take the majority of those submissions, and quickly became the most active subreddits. By the way, it's fascinating that you can still see spikes in gaming subreddits when games are released. This year, the big spikes happened in /r/battlefield3 for its October release, and /r/skyrim for its November release. I should also note here that Minecraft, Starcraft, and TF2 have done a great job of maintaining active communities around their respective video games! Notable subreddits born this year /r/occupywallstreet sprang to life in late 2011, showing that Reddit was still politically active enough to rally around a common cause (despite the waning popularity of /r/politics). Bronies everywhere came together at /r/mylittlepony to rejoice in all that is wonderful about My Little Pony. I've never understood the appeal myself, but it's undeniable that /r/mylittlepony was (and still is) an active part of the Reddit community! Notable subreddits died this year /r/reddit.com went down in a Rasputin-like fashion, coming back to life near the end of 2011 to have yet more posts submitted to it, then finally clunked over the head and put to rest. RIP. 2012 - The Year Picture Subreddits Took Over Finally, here we are in 2012. Removing /r/reddit.com seemed to make pictures even more popular on Reddit. Just check the front page on January 1, 2012 (77/100 top posts are images). At this rate, it looks like Reddit will be an image board within a few years. 2012 - The Year Picture Subreddits Took Over Apparently Redditors got tired of the rage comics in 2012, as shown by the declining activity in /r/f7u12. The Pokemon craze hit Reddit big time in 2012, with /r/pokemon finally making it into the top 25 active subreddits. The most interesting thing to me, though, is how /r/trees remained remarkably active despite the fact that it's not a default subreddit. Reddit has some seriously dedicated tokers! /r/politics experienced a small resurgence of activity near November for the 2012 Presidential elections, but it was nothing compared to last time. If anything, this fact shows how much Reddit has changed in the past four years: from politically active tech nerds to gamers, stoners, and picture/meme lovers. Finally, we saw two more subreddit spikes due to game releases this year: /r/Diablo for Diablo 3's mid-May release, and /r/guildwars2 for its late August release. Notable subreddits born this year /r/ModerationLog was created to log all of the submissions that were removed by moderators of Reddit's various subreddits. There seems to be quite a lot of moderation going on nowadays! /r/POLITIC's post bot has excelled at bringing content into this subreddit, made obvious by the sudden explosion of /r/POLITIC's posts in late 2012 (not shown here). It's pretty rare for a new subreddit to make a dent in Reddit's overall submission traffic nowadays, so hats off to the /r/POLITIC bot! Notable subreddits died this year While a few subreddits seem to have fallen from the charts this year (/r/firstworldproblems, /r/IAmA, /r/worldnews), they're still alive and well. Only time will tell if text posts and article links will be completely replaced by images and videos. One last word on the future of Reddit I have one last comment on the future of Reddit. I've seen a lot of people getting disparaged by how Reddit is slowly becoming a sophisticated image board, and I'd like to address that concern here. Let's recall the subreddit diversity graph I showed earlier: Subreddit diversity over time What's readily apparent from this graph is that even though the picture subreddits are dominating Reddit in terms of pure number of submissions, there are new subreddits dedicated to every topic you can imagine being created every day. Yet the biggest hurdle Redditors face is finding these new subreddits that are interesting to them. Oftentimes a new subreddit springs into popularity because of a chance mention in the comments of a picture submission, but that's not really an efficient way of guiding Redditors to the subreddits that they find interesting. As such, I think Reddit needs a new tool -- or set of tools -- to help new Redditors find smaller subreddits that interest them. What should that new tool be? --- ## Why I love open source, and wish science worked this way URL: https://www.randalolson.com/2013/03/02/why-i-love-open-source-and-wish-science-worked-this-way/ Published: 2013-03-02 Categories: open science, philosophy, research Tags: open science, open source Randy Olson discusses open source and how its philosophies could benefit the scientific process. Yesterday, I got bored and hastily hacked together a script to scrape word frequencies from Reddit and make word clouds out of them. Of course, I included the source code on github so everyone else could use the script if they wanted to. Overnight, some people I've never met before found my repo, decided they loved it, improved my script 1000x, committed their changes to my github repo, then started promoting my github repo through venues I'd never even thought of. Isn't that amazing? Just because I openly posted my code on the internet, a group of people came together to refine my work and make it better for everyone. In the end, I learned some new coding tricks, the code is reaching a broader audience of people that I'd never imagined, and everyone ended up with a better final product. Wouldn't it be nice if science worked this way? --- ## Fun with the Python Reddit API Wrapper and word clouds URL: https://www.randalolson.com/2013/03/01/fun-with-the-python-reddit-api-wrapper-and-word-clouds/ Published: 2013-03-01 Categories: analysis, python, reddit Tags: analysis, reddit, word cloud, word frequencies Randy Olson demonstrates a new Python library that scrapes word frequencies from Reddit and make word clouds with them. I got bored today and threw together some Python code to scrape word frequencies from Reddit and make word clouds. Everyone on Reddit seemed to love them, so I put them up on github so everyone could start making their own word clouds. All that's really left to do is connect these scripts to a word cloud generating library so we don't even have to copy & paste text into Wordle any more. If you're up to the task, please email me and fork away on github. Making word clouds for subreddits is a surprisingly effective way to get a gist for what a subreddit is really talking about. Take /r/evolution, for example. They're serious business about evolution. Others were more amusing. /r/trees, for example, seems to be preoccupied with cursing about things. whereas /r/aww can be concisely described by "upvote cats, fuck humans." Even the /r/space nerds seem to get riled up when discussing NASA, terraforming, and meteorites. Come join in on the fun and make some word clouds for your favorite subreddit: https://github.com/rhiever/reddit-analysis --- ## How do you maximize the Tweetability of your presentations? URL: https://www.randalolson.com/2013/02/16/how-do-you-maximize-the-tweetability-of-your-presentations/ Published: 2013-02-16 Categories: outreach Tags: hashtag, impact, outreach, presentations, reporting, tweet, twitter Given the rising popularity of Twitter reporting at conferences, how do we make our presentations more accessible to Twitter? Given that conference season is coming up, I've been giving a great deal of thought to the best ways to increase the impact of my presentations. How do I make them memorable? How do I make them available and interesting to the widest audience possible? I've especially taken notice of the rising popularity of "Twitter reporters" at scientific conferences, who articulately summarize an entire year or more worth of research by the current presenter into one or two hashtagged 140-character tweets, then provide a link to more details for those interested. For example, lately, I've been remotely following the annual AAAS meeting via the Twitter hashtag #aaasmtg. Thanks to the Twitter reporters, I'm able to follow what's going on at the conference and stay up to date with people who share similar research interests, all from the comfort of my own mobile phone. That's what got me thinking: Twitter reporting is a relatively new thing, and designing your presentations with Twitter reporting in mind is probably even more rare. How can we turn this dark art into a formal process? Here are a few informal recommendations I've run across. Establish a Twitter hashtag for your conference so the Twitter reporters and followers all have the same hashtag to refer to. Start your presentation with a disclaimer stating whether and how your presentation can be shared via Twitter and other forms of social media. Doing so not only makes it clear that you are okay with sharing the information in your presentation, but may also motivate someone in the audience to become a Twitter reporter for your presentation. I attended one of Ethan White's recent talks, and his second and third slides provide a great template for a disclaimer. Make your Twitter and social media handles known to the audience so they can follow you, communicate with you, and most importantly communicate about you. Make your slides publicly available before the presentation and provide a short link to them (e.g., bit.ly) so people can discuss them on Twitter during and after your presentation. FigShare seems to be a popular option for publicly hosting slides, and they even give you a DOI. Clearly use slogans in your presentations that effectively summarize the main points. Something concise and catchy like "Studying the effective collective" like in Iain Couzin's recent swarming work. Do you have any more? Feel free to share and discuss in the comments. --- ## IPython Notebook workshop report: still plagued by installation issues URL: https://www.randalolson.com/2013/01/24/ipython-notebook-workshop-report-still-plagued-by-installation-issues/ Published: 2013-01-24 Categories: ipython, outreach Tags: ipython, notebook, report, workshop A brief report from a workshop run to teach scientists how to use IPython Notebook for scientific computing. Update (Nov. 2014): IPython Notebook installation has advanced considerably since I originally published this post. Check out the Anaconda Python Distribution for an easy one-click installer for all your Python library needs. Today I ran a small (~20 person) 1-hour workshop at Michigan State University focusing on installing IPython Notebook and using it as a research notebook. Since I knew it's quite a pain to install IPython Notebook for the first time, I put together a set of installation instructions for both Windows and Mac/Linux and asked the attendees to attempt to install IPython Notebook ahead of time. On top of that, I held a 30-minute session before the workshop dedicated solely to installing IPython Notebook (which admittedly, helped a little bit). By the end of it, I was able to get through the IPython Notebook demo, and many people seemed excited about using it for their research. Before everyone left, I had the attendees fill out a post-workshop survey to get their opinion on the workshop. In the post-workshop survey, I found some interesting trends in the responses: All but one attendee indicated that they understood how to use IPython Notebook and found its interface intuitive Half of the workshop attendees indicated that they planned to start using IPython Notebook for their research Those that indicated that they don't plan to use IPython Notebook for their research cited installation issues as the primary or secondary reason for not doing so I think the third trend is the most telling: one of the biggest barriers to the adoption of IPython Notebook is installation issues. This isn't an uncommon observation, either. In my case, despite all of my precautions, the workshop was still plagued by installation issues. At least 1/4 of the workshop attendees had these issues during the workshop, including: Missing libraries being required for the newer version of IPython Notebook that don't come with the installation package (not even EPD) by default IPython Notebook not running in Chrome, only Firefox pandas compatibility issues with the newer version IPython Notebook Newer ipynb files not being backwards-compatible with older IPython Notebook versions This highlights two key some key issues the IPython community needs to work on before IPython can reach a broader audience: (1) we cannot rely on EPD to maintain a free, update-to-date Python package installer and (2) we need to sort out the IPython installation process and distill it to a single, double-click install file. UPDATE: Fernando Perez offered some easy installation instructions that seem to have solved the installation issues. I will report on how these work out in an upcoming workshop soon! --- ## Filling in Python's gaps in statistics packages with Rmagic URL: https://www.randalolson.com/2013/01/14/filling-in-pythons-gaps-in-statistics-packages-with-rmagic/ Published: 2013-01-14 Categories: ipython, productivity, python, statistics, tutorial Tags: ipython, pandas, python, R language, Rmagic, statistics, tutorial Randy Olson shows how to use IPython's Rmagic package to seamlessly integrate R code with the IPython interface. Have you ever found yourself searching for a statistics package in Python, but it just isn't available? This is the biggest reason I've heard when my colleagues say they're unwilling to make the switch from R to Python for statistical analysis. To counteract that argument, the geniuses developing IPython created Rmagic, a package which allows you to run R code within the IPython interface. Let's get to the cut and dry and see how it works. If you want to follow along, the IPython Notebook and data file is available on github: https://github.com/rhiever/rmagic-tutorial Installing Rmagic Rmagic comes along with the IPython package, so just follow my IPython tutorial to install IPython. Beyond IPython, all you need is Python's rpy2 package to run Rmagic: sudo easy_install rpy2 I also exclusively use the pandas package when I'm dealing with data, so make sure to have that installed as well: sudo easy_install pandas Using Rmagic for statistical analysis I do most of my statistical analysis in Python nowadays, but sometimes there's just that one statistical function that I need that Python doesn't quite have yet. In those cases, I use Rmagic to "fill the gap" in Python's statistics packages. Rmagic lets me pass my data to R, run the R function on the data, then seamlessly return the data back to Python before I start having nightmares about using R again. Since there are already a ton of R statistics tutorials out there, this tutorial will instead concentrate on how to use Rmagic to link together Python and R code. Load the rmagic extension %load_ext rmagic Use R to view summary information about the data If you like how R provides summary information about the data, you can print out the summary() function. %R tells IPython that the line is R code. -i parasiteData tells IPython to pass the parasiteData DataFrame to R. The rest is R code. You can have multiple lines of R code separated by semicolons. Note that when you want to print anything out from R, you need to place the print() function around it. from pandas import * # read data from data file into a pandas DataFrame parasiteData = read_csv("parasite_data.csv", sep=",", na_values=["", " "]) # print the R summary() function for the data %R -i parasiteData print(summary(parasiteData)) Virulence Replicate ShannonDiversity Min. :0.50 Min. : 1.0 Min. :0.0000 1st Qu.:0.60 1st Qu.:13.0 1st Qu.:0.0000 Median :0.75 Median :25.5 Median :0.8457 Mean :0.75 Mean :25.5 Mean :0.8364 3rd Qu.:0.90 3rd Qu.:38.0 3rd Qu.:1.5337 Max. :1.00 Max. :50.0 Max. :2.9008 NA's :50 View a R plot in IPython You can even view plots of your data in IPython from R. from pandas import * # read data from data file into a pandas DataFrame parasiteData = read_csv("parasite_data.csv", sep=",", na_values=["", " "]) # plot the Shannon Diversity as a function of Virulence in R %R -i parasiteData plot(ShannonDiversity ~ Replicate, data = parasiteData, xlab="Replicate", ylab="Shannon Diversity") Filling in the gaps Beyond the basics, you can use R to perform a single function on your data, then return the result back to Python. Here's a simple example. -o meanShannonDiversity tells IPython to return the meanShannonDiversity variable back from R to IPython. from pandas import * # read data from data file into a pandas DataFrame parasiteData = read_csv("parasite_data.csv", sep=",", na_values=["", " "]) # calculate the mean Shannon Diversity for experiments with Virulence = 0.7 # do the subsetting in Python - it's easier! ShannonDiversity = parasiteData[parasiteData["Virulence"] == 0.7]["ShannonDiversity"] %R -i ShannonDiversity -o meanShannonDiversity meanShannonDiversity How about a more useful example? We can use R's built-in bootstrapping function to construct 95% confidence intervals for our data. from pandas import * # read data from data file into a pandas DataFrame parasiteData = read_csv("parasite_data.csv", sep=",", na_values=["", " "]) # calculate the 95% confidence interval for Shannon Diversity for # experiments with Virulence = 0.5 and 0.8 # do the subsetting in Python - it's easier! ShannonDiversityV5 = parasiteData[parasiteData["Virulence"] == 0.5]["ShannonDiversity"] ShannonDiversityV8 = parasiteData[parasiteData["Virulence"] == 0.8]["ShannonDiversity"] # load R's boot library %R require(boot) # define the function we're bootstrapping (mean) %R sampleMean And there you have it: we know there is a significant difference between the experiments with Virulence = 0.5 and 0.8 because their 95% confidence intervals don't overlap. For those who use Octave, there is a similar octavemagic package in IPython as well. --- ## Neuroevolution: an alternative route to Artificial Intelligence URL: https://www.randalolson.com/2012/08/13/neuroevolution-an-alternative-route-to-artificial-intelligence/ Published: 2012-08-13 Categories: philosophy, research Tags: artificial brain, artificial intelligence, evolution, neurobiology, neuroevolution, neuroscience, strong ai Randy Olson discusses neuroevolution as an alternative route to Artificial Intelligence. If you were to ask a random person what the best example of Artificial Intelligence is out there, what do you think it would be? Most likely, it would be IBM's Watson. In a stunning display of knowledge and accuracy, Watson blew away the world Jeopardy champions Ken Jennings and Brad Rutter without blowing a fuse, and ended with Jennings proclaiming, "I for one welcome our new computer overlords." IBM's Watson represents the current popular approach to AI: that is, spending hundreds of hours hand-coding and fine-tuning a program to perform exceedingly well on a single task. Most people in the field of AI call machines like Watson an expert system because they are designed to be experts at a single task. This approach has been wildly successful lately, producing machines that drive cars and fly UAVs by themselves, beat world chess and Jeopardy champions, and even fool some people into thinking they're human. However, imagine how hard it would be to hand-code a system that could do everything the human brain is capable of. Do you think that sounds impossible? That's the reason why the field of neuroevolution was born: scientists wanted to harness the creative power of evolution to design the programs that could achieve human-level intelligence. What is Neuroevolution? Neuroevolution, or neuro-evolution, is a form of machine learning that uses evolutionary algorithms to train artificial neural networks. It is useful for applications such as games and robot motor control, where it is easy to measure a network's performance at a task but difficult or impossible to create a syllabus of correct input-output pairs for use with a supervised learning algorithm. -Wikipedia What does all that mean? Broadly speaking, the goal of neuroevolution is to evolve an artificial brain with a genetic algorithm to solve a specific task. The artificial brain, oftentimes called the artificial neural network, is designed based off of our understanding of how biological brains work. This video does a great job of explaining artificial neural networks: As the video mentioned, oftentimes the genetic algorithm starts out with a bunch of random artificial brains. The genetic algorithm then emulates the process of evolution: Fitness evaluation: each of the artificial brains are tested on how well they perform at a task. Selection: the brains that perform better are chosen to reproduce into the next generation of artificial brains. Descent with modification: the offspring of those artificial brains are created as copies of their parent brains with slight modifications. This process repeats over and over until the artificial brains master the task. Here's an example of an artificial brain being evolved to walk in a two-legged robot. Notice how the artificial brain does a really bad job of walking at first, but eventually learns walk without falling at all. Why is that useful? Genetic algorithms have been proven to be a creative and powerful designer. For example, researchers once used a genetic algorithm to design an antenna for one of NASA's satellites. The original antenna took months for engineers to design; cost thousands of dollars per antenna; and didn't even perform as well as NASA had hoped. An entrepreneurial group of researchers at UCSC decided to make an attempt at designing their own version of the antenna with a genetic algorithm, and evolved an antenna that used a single piece of wire that cost next to nothing and performed better than the antenna designed by the engineers. The same concept applies for evolving artificial brains. Researchers at UT Austin have evolved artificial brains to control a rocket into space without fins, which is an otherwise extremely difficult problem to engineer. [videos] Meanwhile, researchers at UCF have evolved artificial brain controllers for two-legged robots that walk and balance all by themselves. [video] Evolved artificial brains are even being used in video games, such as UT Austin's NERO video game. [video] There are plenty more examples of "neuroevolution in action" out there; these are just a few choice examples. Neuroevolution has a promising future of designing intelligent algorithms for robot control, vehicle navigation, and many, many, many more applications. Neuroevolution and Artificial Intelligence The real advantage of neuroevolution is what it brings to the development of Artificial Intelligence. In the past, computer scientists working on AI would design an algorithm that would exhibit intelligent behavior, then tweak that algorithm's parameters until it exhibited "optimal" intelligent behavior. The AI they designed either worked or it didn't, and oftentimes their results didn't teach us much about how human brains work. On the other hand, in neuroevolution, scientists can begin to ask questions about the evolution of human-level intelligence: "What challenges (or set of challenges) were ancient organisms faced with that required them to evolve intelligence to succeed?" "What were the ‘building blocks’ to human-level intelligence?" etc. Indeed, neuroevolution promises to be an insightful field of study, since scientists can not only attempt to create an artificial intelligence, but also hypothesize about how intelligence was created in the first place. (Which is why neuroscientists and biologists are also interested and involved in this field!) --- ## Insight from "Don't be such a scientist" URL: https://www.randalolson.com/2012/08/10/insight-from-dont-be-such-a-scientist/ Published: 2012-08-10 Categories: outreach Tags: communication, don't be such a scientist, outreach, randy olson, science Randy Olson discusses his insight from the other Prof. Randy Olson's book, "Don't be such a scientist." This week, I sat down to read Randy Olson's book, "Don't be such a scientist." Other than the fact that Randy and I share the same name (and nearly the same profession!... or at least, we used to), we apparently share the same interest of effectively communicating science to the public. When I finished reading his book, I started reading the reviews online, and was sorely disappointed with one common problem among them: the main message of Randy's book seems to have been lost to most everyone! Many reviewers critique his book as if it were a guide on how to communicate science to the public, or especially on how to spice up your powerpoints so people won't doze off by slide 4. If you watched Randy's documentary Flock of Dodos (which is an instant watch on Netflix, by the way), it's clear his mission goes beyond writing a self-help book on "communicating your science to the public." The United States of America is facing a serious epidemic right now: nearly half the nation, ~150 million people, don't think Darwinian evolution is a reality, despite the overwhelming evidence supporting it. The Intelligent Design movement, with its catchy slogan "Teach the controversy" and millions of dollars being poured into public relations, is clearly winning this battle at present, despite the fact that they have no data whatsoever legitimately supporting their position. I don't think I have to belabor the point of what an impact it would have on society if the majority of the U.S. public distrusted scientific reasoning. What's worse, some of the greatest scientific minds of our time don't seem to care that we're losing this battle. This is where I think Randy's book really shines: he offers a new strategy to combat the anti-science movement. The attackers of science are a potential communication opportunity. They are a source of tension and conflict. They can actually be used to tell a more interesting story, one that can grab the interest of a much wider audience. We shouldn't look at the anti-science movement as an annoyance, or a group of idiots to disdainfully look down upon. Instead, we should see them as an opportunity. An opportunity to tell a story that the public is longing to hear: where scientists are the superheros saving their innocent mothers from the lies and persuasions of the treacherous, power-hungry men from the evil anti-science empire. The public eats this stuff up; why do you think all the popular movies follow the same, general plot line every time? As scientists, we've been afforded a rare opportunity to bring ourselves into the light of the public as the "knights of truth and justice" that we are. Why aren't we taking advantage of it? Randy's book does a much better job of explaining this idea and why it's so promising, and Flock of Dodos stands as a shining example of his idea in action. Check them out, then come join your fellow Knights of Truth and Justice in their crusade against the evil Anti-Science Empire. --- ## Statistical analysis made easy in Python with SciPy and pandas DataFrames URL: https://www.randalolson.com/2012/08/06/statistical-analysis-made-easy-in-python/ Published: 2012-08-06 Categories: ipython, productivity, python, statistics, tutorial Tags: analysis of variance, ANOVA, bootstrap, confidence interval, data management, ipython, Mann-Whitney-Wilcoxon, MWW, notebook, pandas, plotting data, python, RankSum, research, standard error, statistics, tutorial Randy Olson demonstrates how to use SciPy and pandas DataFrames to perform commonly-used statistical analyses and tests in Python. I finally got around to finishing up this tutorial on how to use pandas DataFrames and SciPy together to handle any and all of your statistical needs in Python. This is basically an amalgamation of my two previous blog posts on pandas and SciPy. This is all coded up in an IPython Notebook, so if you want to try things out for yourself, everything you need is available on github: https://github.com/briandconnelly/BEACONToolkit/tree/master/analysis/scripts Statistical Analysis in Python In this section, we introduce a few useful methods for analyzing your data in Python. Namely, we cover how to compute the mean, variance, and standard error of a data set. For more advanced statistical analysis, we cover how to perform a Mann-Whitney-Wilcoxon (MWW) RankSum test, how to perform an Analysis of variance (ANOVA) between multiple data sets, and how to compute bootstrapped 95% confidence intervals for non-normally distributed data sets. Python's SciPy Module The majority of data analysis in Python can be performed with the SciPy module. SciPy provides a plethora of statistical functions and tests that will handle the majority of your analytical needs. If we don't cover a statistical function or test that you require for your research, SciPy's full statistical library is described in detail at: http://docs.scipy.org/doc/scipy/reference/tutorial/stats.html Python's pandas Module The pandas module provides powerful, efficient, R-like DataFrame objects capable of calculating statistics en masse on the entire DataFrame. DataFrames are useful for when you need to compute statistics over multiple replicate runs. For the purposes of this tutorial, we will use Luis Zaman's digital parasite data set: from pandas import * # must specify that blank space " " is NaN experimentDF = read_csv("parasite_data.csv", na_values=[" "]) print experimentDF [class 'pandas.core.frame.DataFrame'] Int64Index: 350 entries, 0 to 349 Data columns: Virulence 300 non-null values Replicate 350 non-null values ShannonDiversity 350 non-null values dtypes: float64(2), int64(1) Accessing data in pandas DataFrames You can directly access any column and row by indexing the DataFrame. # show all entries in the Virulence column print experimentDF["Virulence"] 0 0.5 1 0.5 2 0.5 3 0.5 4 0.5 ... 346 NaN 347 NaN 348 NaN 349 NaN Name: Virulence, Length: 350 # show the 12th row in the ShannonDiversity column print experimentDF["ShannonDiversity"][12] 1.58981 You can also access all of the values in a column meeting a certain criteria. # show all entries in the ShannonDiversity column > 2.0 print experimentDF[experimentDF["ShannonDiversity"] > 2.0] Virulence Replicate ShannonDiversity 8 0.5 9 2.04768 89 0.6 40 2.01066 92 0.6 43 2.90081 96 0.6 47 2.02915 ... 235 0.9 36 2.19565 237 0.9 38 2.16535 243 0.9 44 2.17578 251 1.0 2 2.16044 Blank/omitted data (NA or NaN) in pandas DataFrames Blank/omitted data is a piece of cake to handle in pandas. Here's an example data set with NA/NaN values. import numpy as np print experimentDF[np.isnan(experimentDF["Virulence"])] Virulence Replicate ShannonDiversity 300 NaN 1 0.000000 301 NaN 2 0.000000 302 NaN 3 0.833645 303 NaN 4 0.000000 ... 346 NaN 47 0.000000 347 NaN 48 0.444463 348 NaN 49 0.383512 349 NaN 50 0.511329 DataFrame methods automatically ignore NA/NaN values. print "Mean virulence across all treatments:", experimentDF["Virulence"].mean() Mean virulence across all treatments: 0.75 However, not all methods in Python are guaranteed to handle NA/NaN values properly. from scipy import stats print "Mean virulence across all treatments:", stats.sem(experimentDF["Virulence"]) Mean virulence across all treatments: nan Thus, it behooves you to take care of the NA/NaN values before performing your analysis. You can either: (1) filter out all of the entries with NA/NaN # NOTE: this drops the entire row if any of its entries are NA/NaN! print experimentDF.dropna() [class 'pandas.core.frame.DataFrame'] Int64Index: 300 entries, 0 to 299 Data columns: Virulence 300 non-null values Replicate 300 non-null values ShannonDiversity 300 non-null values dtypes: float64(2), int64(1) If you only care about NA/NaN values in a specific column, you can specify the column name first. print experimentDF["Virulence"].dropna() 0 0.5 1 0.5 2 0.5 3 0.5 ... 296 1 297 1 298 1 299 1 Name: Virulence, Length: 300 (2) replace all of the NA/NaN entries with a valid value print experimentDF.fillna(0.0)["Virulence"] 0 0.5 1 0.5 2 0.5 3 0.5 4 0.5 ... 346 0 347 0 348 0 349 0 Name: Virulence, Length: 350 Take care when deciding what to do with NA/NaN entries. It can have a significant impact on your results! print ("Mean virulence across all treatments w/ dropped NaN:", experimentDF["Virulence"].dropna().mean()) print ("Mean virulence across all treatments w/ filled NaN:", experimentDF.fillna(0.0)["Virulence"].mean()) Mean virulence across all treatments w/ dropped NaN: 0.75 Mean virulence across all treatments w/ filled NaN: 0.642857142857 Mean of a data set The mean performance of an experiment gives a good idea of how the experiment will turn out on average under a given treatment. Conveniently, DataFrames have all kinds of built-in functions to perform standard operations on them en masse: `add()`, `sub()`, `mul()`, `div()`, `mean()`, `std()`, etc. The full list is located at: http://pandas.pydata.org/pandas-docs/stable/api.html#computations-descriptive-stats Thus, computing the mean of a DataFrame only takes one line of code: from pandas import * print ("Mean Shannon Diversity w/ 0.8 Parasite Virulence =", experimentDF[experimentDF["Virulence"] == 0.8]["ShannonDiversity"].mean()) Mean Shannon Diversity w/ 0.8 Parasite Virulence = 1.2691338188 Variance in a data set The variance in the performance provides a measurement of how consistent the results of an experiment are. The lower the variance, the more consistent the results are, and vice versa. Computing the variance is also built in to pandas DataFrames: from pandas import * print ("Variance in Shannon Diversity w/ 0.8 Parasite Virulence =", experimentDF[experimentDF["Virulence"] == 0.8]["ShannonDiversity"].var()) Variance in Shannon Diversity w/ 0.8 Parasite Virulence = 0.611038433313 Standard Error of the Mean (SEM) Combined with the mean, the SEM enables you to establish a range around a mean that the majority of any future replicate experiments will most likely fall within. pandas DataFrames don't have methods like SEM built in, but since DataFrame rows/columns are treated as lists, you can use any NumPy/SciPy method you like on them. from pandas import * from scipy import stats print ("SEM of Shannon Diversity w/ 0.8 Parasite Virulence =", stats.sem(experimentDF[experimentDF["Virulence"] == 0.8]["ShannonDiversity"])) SEM of Shannon Diversity w/ 0.8 Parasite Virulence = 0.110547585529 A single SEM will usually envelop 68% of the possible replicate means and two SEMs envelop 95% of the possible replicate means. Two SEMs are called the "estimated 95% confidence interval." The confidence interval is estimated because the exact width depend on how many replicates you have; this approximation is good when you have more than 20 replicates. Mann-Whitney-Wilcoxon (MWW) RankSum test The MWW RankSum test is a useful test to determine if two distributions are significantly different or not. Unlike the t-test, the RankSum test does not assume that the data are normally distributed, potentially providing a more accurate assessment of the data sets. As an example, let's say we want to determine if the results of the two following treatments significantly differ or not: # select two treatment data sets from the parasite data treatment1 = experimentDF[experimentDF["Virulence"] == 0.5]["ShannonDiversity"] treatment2 = experimentDF[experimentDF["Virulence"] == 0.8]["ShannonDiversity"] print "Data set 1:\n", treatment1 print "Data set 2:\n", treatment2 Data set 1: 0 0.059262 1 1.093600 2 1.139390 3 0.547651 ... 45 1.937930 46 1.284150 47 1.651680 48 0.000000 49 0.000000 Name: ShannonDiversity Data set 2: 150 1.433800 151 2.079700 152 0.892139 153 2.384740 ... 196 2.077180 197 1.566410 198 0.000000 199 1.990900 Name: ShannonDiversity A RankSum test will provide a P value indicating whether or not the two distributions are the same. from scipy import stats z_stat, p_val = stats.ranksums(treatment1, treatment2) print "MWW RankSum P for treatments 1 and 2 =", p_val MWW RankSum P for treatments 1 and 2 = 0.000983355902735 If P If the treatments do not significantly differ, we could expect a result such as the following: treatment3 = experimentDF[experimentDF["Virulence"] == 0.8]["ShannonDiversity"] treatment4 = experimentDF[experimentDF["Virulence"] == 0.9]["ShannonDiversity"] print "Data set 3:\n", treatment3 print "Data set 4:\n", treatment4 Data set 3: 150 1.433800 151 2.079700 152 0.892139 153 2.384740 ... 196 2.077180 197 1.566410 198 0.000000 199 1.990900 Name: ShannonDiversity Data set 4: 200 1.036930 201 0.938018 202 0.995956 203 1.006970 ... 246 1.564330 247 1.870380 248 1.262280 249 0.000000 Name: ShannonDiversity # compute RankSum P value z_stat, p_val = stats.ranksums(treatment3, treatment4) print "MWW RankSum P for treatments 3 and 4 =", p_val MWW RankSum P for treatments 3 and 4 = 0.994499571124 With P > 0.05, we must say that the distributions do not significantly differ. Thus changing the parasite virulence between 0.8 and 0.9 does not result in a significant change in Shannon Diversity. One-way analysis of variance (ANOVA) If you need to compare more than two data sets at a time, an ANOVA is your best bet. For example, we have the results from three experiments with overlapping 95% confidence intervals, and we want to confirm that the results for all three experiments are not significantly different. treatment1 = experimentDF[experimentDF["Virulence"] == 0.7]["ShannonDiversity"] treatment2 = experimentDF[experimentDF["Virulence"] == 0.8]["ShannonDiversity"] treatment3 = experimentDF[experimentDF["Virulence"] == 0.9]["ShannonDiversity"] print "Data set 1:\n", treatment1 print "Data set 2:\n", treatment2 print "Data set 3:\n", treatment3 Data set 1: 100 1.595440 101 1.419730 102 0.000000 103 0.000000 ... 146 0.000000 147 1.139100 148 2.383260 149 0.056819 Name: ShannonDiversity Data set 2: 150 1.433800 151 2.079700 152 0.892139 153 2.384740 ... 196 2.077180 197 1.566410 198 0.000000 199 1.990900 Name: ShannonDiversity Data set 3: 200 1.036930 201 0.938018 202 0.995956 203 1.006970 ... 246 1.564330 247 1.870380 248 1.262280 249 0.000000 Name: ShannonDiversity # compute one-way ANOVA P value from scipy import stats f_val, p_val = stats.f_oneway(treatment1, treatment2, treatment3) print "One-way ANOVA P =", p_val One-way ANOVA P = 0.381509481874 If P > 0.05, we can claim with high confidence that the means of the results of all three experiments are not significantly different. Bootstrapped 95% confidence intervals Oftentimes in wet lab research, it's difficult to perform the 20 replicate runs recommended for computing reliable confidence intervals with SEM. In this case, bootstrapping the confidence intervals is a much more accurate method of determining the 95% confidence interval around your experiment's mean performance. Unfortunately, SciPy doesn't have bootstrapping built into its standard library yet. However, there is already a scikit out there for bootstrapping. Enter the following command to install it: sudo easy_install scikits.bootstrap Bootstrapping 95% confidence intervals around the mean with this function is simple: # subset a list of 10 data points treatment1 = experimentDF[experimentDF["Virulence"] == 0.8]["ShannonDiversity"][:10] print "Small data set:\n", treatment1 Small data set: 150 1.433800 151 2.079700 152 0.892139 153 2.384740 154 0.006980 155 1.971760 156 0.000000 157 1.428470 158 1.715950 159 0.000000 Name: ShannonDiversity import scipy import scikits.bootstrap as bootstrap # compute 95% confidence intervals around the mean CIs = bootstrap.ci(data=treatment1, statfunction=scipy.mean) print "Bootstrapped 95% confidence intervals\nLow:", CIs[0], "\nHigh:", CIs[1] Bootstrapped 95% confidence intervals Low: 0.659028048 High: 1.722468024 Note that you can change the range of the confidence interval by setting the alpha: # 80% confidence interval CIs = bootstrap.ci(treatment1, scipy.mean, alpha=0.2) print "Bootstrapped 80% confidence interval\nLow:", CIs[0], "\nHigh:", CIs[1] Bootstrapped 80% confidence interval Low: 0.827291024 High: 1.5420059 And also modify the size of the bootstrapped sample pool that the confidence intervals are taken from: # bootstrap 20,000 samples instead of only 10,000 CIs = bootstrap.ci(treatment1, scipy.mean, n_samples=20000) print ("Bootstrapped 95% confidence interval w/ 20,000 samples\nLow:", CIs[0], "\nHigh:", CIs[1]) Bootstrapped 95% confidence interval w/ 20,000 samples Low: 0.644756972 High: 1.7071459 Generally, bootstrapped 95% confidence intervals provide more accurate confidence intervals than 95% confidence intervals estimated from the SEM. --- ## David Eagleman: are we taking the right approach to Artificial Intelligence? URL: https://www.randalolson.com/2012/06/28/david-eagleman-are-we-taking-the-right-approach-to-artificial-intelligence/ Published: 2012-06-28 Categories: philosophy, research Tags: artificial brains, artificial intelligence, neuroevolution, philosophy Randy Olson ponders David Eagleman's suggestion that we are taking the wrong approach to Artificial Intelligence. I watched a YouTube video earlier today of an interview with David Eagleman, where he discussed his thoughts on the current approach that most researchers are taking to the problem of Artificial Intelligence. To me, this is an extremely interesting topic to ponder. He put to words a good portion of what has been on my mind about the field of AI. I believe Eagleman is on the right track. Let's start by looking up the definition of intelligence: intelligence 1. capacity for learning, reasoning, understanding, and similar forms of mental activity; aptitude in grasping truths, relationships, facts, meanings, etc. Ignoring hard-coded solutions that exhibit intelligent behavior (because that is hard-coded pseudo intelligence; I don't even consider it AI), some AI techniques (machine learning, NLP) have shown the ability to learn and understand the relationships between things. But can they reason about those facts? Can they understand the true meaning of those facts? Further, can they reason about those facts to the point where they can create new ideas without being taught them? As far as I know, the answer to all three is those questions is "no." (Please point me to the papers if you have found otherwise! I'd be very interested.) Why? I highly doubt it's because the people working on these problems are stupid. It's more likely because the approach we've been taking is the wrong approach. As was mentioned in the video, the brain doesn't solve problems by solving a sub-problem for every task and combining them together. One solution can also be the solution to an entirely different problem, or an amalgamation of the subsets of two solutions can be the solution to another problem. Breaking problems down into sub-problems and solving them individually is an inefficient approach when it comes to creating a true artificial general intelligence. Most importantly, here's something to consider: what is the only method by which we've seen intelligence be created on Earth? It wasn't made by man nor by another intelligence; it was made by evolution over extremely long periods of time. Why, then, do we ignore this fact and set aside one of the most powerful creative tools we have available to us? Below, I responded to a few criticisms of his video: It's very easy to say (just a hypothetical quote, not Eagleman's) "No, no, we're approaching this all wrong. We can't do this in a piecemeal fashion. We need to approach this holistically." That's great and all... but how do you propose we do that in terms that are concrete enough that we can actually act on them (sadly, "watching how nature does it" is not enough)? There are entire sub-fields of AI dedicated to this. Neuroevolution(ists?), for example, evolve artificial neural networks (ANNs, or "artificial brains," if you will) with the task of solving a specific task. A fairly recent advent in this field is the multi-objective evolution of those ANNs, whereby the ANNs are evolved to solve a set of tasks. From there, we can design experiments that ask, "What task challenges (or set of task challenges) were ancient organisms faced with that required them to evolve intelligence to succeed?," "What were the 'building blocks' to intelligence?," etc. Indeed, this field promises to be extremely insightful, since we can not only attempt to create a general AI, but also hypothesize about how general intelligence was created in the first place. (Which is why neuroscientists are also involved in this field.) Furthermore, are you absolutely certain that all of the subproblems we are solving won't aid in building that system? No, it's impossible to prove that anything won't ever do something. Sure, we could probably create a general AI if we kept at it like this for another 1,000 years or something. (Just think about how many sub-problems human brains have to solve!) It's more a question of: which approach do we think is more fruitful? Should we continue following the approach that has accomplished relatively little in the past ~60 years1, or should we try a new approach that has a much more solid philosophical grounding? 1 Admit it: it's ridiculous that the best AI has to offer are expensive machines that can drive a car or play Jeopardy/chess after ~60 years. Teenagers learn how to drive cars, and any trained person could win at Jeopardy if they had a similarly-tailored database of information that Watson had. I mean no disrespect towards David Eagleman, but I wonder how much he knows about programming. Or specifically, modeling. I'm reasonably certain he got most of his facts right. But the reason we don't have AI yet is not due to programmers methods. It is because our best computers cannot handle the load that a useful AI would take. Indeed, I believe the true problem in this approach to AI right now is finding the proper way to design an artificial brain. Artificial neural networks? Markov brains? Something else? With our current computational technology, we would likely need a highly distributed system to handle the computation required to simulate an artificial brain. But then, that begs the question: why do we need such powerful hardware to emulate the low-power, (relatively) small-sized computing center within our head? Are we modeling the brain correctly? --- ## Using pandas DataFrames to process data from multiple replicate runs in Python URL: https://www.randalolson.com/2012/06/26/using-pandas-dataframes/ Published: 2012-06-26 Categories: python, statistics, tutorial Tags: data management, pandas, plotting data, python, research, statistics Randy Olson demonstrates how to use pandas DataFrames to process data from multiple replicate runs in Python. Per a recommendation in my previous blog post, I decided to follow up and write a short how-to on how to use pandas to process data from multiple replicate runs in Python. If you do research like mine, you'll often find yourself with multiple datasets from an experiment that you've run in replicate multiple times. There are plenty of ways to manage and process data nowadays, but I've never seen it made so easy as it is with pandas. Installing pandas If you don't already have pandas installed, download it at: https://pandas.pydata.org/getting_started.html Then do the typical python package install process. cd unzipped-pandas-folder-name/ python setup.py build_ext --inplace sudo python setup.py install Note: if you're not using pandas in IPython Notebook, the build_ext --inplace part is unnecessary. Using pandas Below, I'll show you the 23 lines of Python code that I use to read in, process, and plot all of the data from my experiments. After that, I'll break the code block down line-by-line and explain what's happening. from pandas import * import glob dataLists = {} # read data for folder in glob.glob("experiment-data-directory/*"): dataLists[folder.split("/")[1]] = [] for datafile in glob.glob(folder + "/*.csv"): dataLists[folder.split("/")[1]].append(read_csv(datafile)) # calculate stats for data meanDFs = {} stderrDFs = {} for key in dataLists.keys(): keyDF = (concat(dataLists[key], axis=1, keys=range(len(dataLists[key]))) .swaplevel(0, 1, axis=1) .sortlevel(axis=1) .groupby(level=0, axis=1)) meanDFs[key] = keyDF.mean() stderrDFs[key] = keyDF.std().div(sqrt(len(dataLists[key]))).mul(2.0) keyDF = None # plot data for column in meanDFs[key].columns: # don't plot generation over generation - that's pointless! if not (column == "generation"): figure(figsize=(20, 15)) title(column.replace("_", " ").title()) ylabel(column.replace("_", " ").title()) xlabel("Generation") for key in meanDFs.keys(): errorbar(x=meanDFs[key]["generation"], y=meanDFs[key][column], yerr=stderrDFs[key][column], label=key) legend(loc=2) Here's one graph from the end product. Required packages Along with the pandas package, the glob package is extremely useful for aggregating folders and files into a single list so they can be iterated over. from pandas import * import glob Reading data with pandas glob.glob(str) aggregates all of the files and folders matching a given *nix directory expression. for folder in glob.glob("experiment-data-directory/*"): For example, say experiment-data-directory contains 4 other directories: treatment1, treatment2, treatment3, and treatment 4. It will return a list of the directories in string format. print glob.glob("experiment-data-directory/*") >>> ["treatment1", "treatment2", "treatment3", "treatment 4"] Similarly, glob.glob(folder + "/*.csv") will return a list of all .csv files in the given directory. print glob.glob("treatment1/*.csv") >>> ["treatment1/run1.csv", "treatment1/run2.csv", "treatment1/run3.csv"] Finally, line 13 stores all of the pandas DataFrames read in by the pandas read_csv(str) function. read_csv(str) is a powerful function that will take care of reading and parsing your csv files into DataFrames. Make sure to have your column titles at the top of each csv file! dataLists[folder.split("/")[1]].append(read_csv(datafile)) Thus, dataLists maps "treatment1", "treatment2", "treatment3", "treatment4" to their corresponding list of DataFrames, with each DataFrame containing the data of a single run. More on how powerful DataFrames are below! Statistics with pandas This bit of code iterates over each treatment. meanDFs = {} stderrDFs = {} for key in dataLists.keys(): And here's where we see the real power of pandas DataFrames. Line 21 merges the list of DataFrames into a single DataFrame containing every run's data for that treatment. keyDF = (concat(dataLists[key], axis=1, keys=range(len(dataLists[key]))) Line 22 makes it so the run data is grouped on a per-data-column basis instead of a per-run basis. .swaplevel(0, 1, axis=1) Line 23 sorts the column names in ascending order. This is purely for aesthetic purposes. .sortlevel(axis=1) Lastly, line 24 groups all of the replicate run data together by column. .groupby(level=0, axis=1)) Here's an example of how this works in practice: In [12]: x Out[12]: A B C 0 -0.264438 -1.026059 -0.619500 1 0.927272 0.302904 -0.032399 2 -0.264273 -0.386314 -0.217601 3 -0.871858 -0.348382 1.100491 In [13]: y Out[13]: A B C 0 1.923135 0.135355 -0.285491 1 -0.208940 0.642432 -0.764902 2 1.477419 -1.659804 -0.431375 3 -1.191664 0.152576 0.935773 In [14]: glued = pd.concat([x, y], axis=1, keys=['x', 'y']) In [15]: glued Out[15]: x y A B C A B C 0 -0.264438 -1.026059 -0.619500 1.923135 0.135355 -0.285491 1 0.927272 0.302904 -0.032399 -0.208940 0.642432 -0.764902 2 -0.264273 -0.386314 -0.217601 1.477419 -1.659804 -0.431375 3 -0.871858 -0.348382 1.100491 -1.191664 0.152576 0.935773 In [16]: glued.swaplevel(0, 1, axis=1).sortlevel(axis=1) Out[16]: A B C x y x y x y 0 -0.264438 1.923135 -1.026059 0.135355 -0.619500 -0.285491 1 0.927272 -0.208940 0.302904 0.642432 -0.032399 -0.764902 2 -0.264273 1.477419 -0.386314 -1.659804 -0.217601 -0.431375 3 -0.871858 -1.191664 -0.348382 0.152576 1.100491 0.935773 In [17]: glued = glued.swaplevel(0, 1, axis=1).sortlevel(axis=1) In [18]: glued Out[18]: A B C x y x y x y 0 -0.264438 1.923135 -1.026059 0.135355 -0.619500 -0.285491 1 0.927272 -0.208940 0.302904 0.642432 -0.032399 -0.764902 2 -0.264273 1.477419 -0.386314 -1.659804 -0.217601 -0.431375 3 -0.871858 -1.191664 -0.348382 0.152576 1.100491 0.935773 By storing the data this way in pandas DataFrames, you can do all kinds of powerful operations on the data on a per-DataFrame basis. In lines 26 and 27, I compute the mean and standard error of the mean of every column (over an arbitrary number of replicates) for every treatment with just a couple lines. meanDFs[key] = keyDF.mean() stderrDFs[key] = keyDF.std().div(sqrt(len(dataLists[key]))).mul(2.0) keyDF = None DataFrames have all kinds of built-in functions to perform standard operations on them en masse: add(), sub(), mul(), div(), mean(), std(), etc. The full list is located at: http://pandas.pydata.org/pandas-docs/dev/generated/pandas.DataFrame.html Plotting pandas data with matplotlib The code below assumes you have a "generation" column that your data is plotted over. If you use another x-axis, it is easy enough to replace "generation" with whatever you named your x-axis. for column in meanDFs[key].columns: # don't plot generation over generation - that's pointless! if not (column == "generation"): figure(figsize=(20, 15)) title(column.replace("_", " ").title()) ylabel(column.replace("_", " ").title()) xlabel("Generation") for key in meanDFs.keys(): errorbar(x=meanDFs[key]["generation"], y=meanDFs[key][column], yerr=stderrDFs[key][column], label=key) legend(loc=2) You can access each column individually by indexing it with the name of the column you want, e.g. dataframe["column_name"]. Since DataFrames store each column's data as a list, it doesn't even take any extra work to pass the data to matplotlib to plot it. --- ## A short demo on how to use IPython Notebook as a research notebook URL: https://www.randalolson.com/2012/05/12/a-short-demo-on-how-to-use-ipython-notebook-as-a-research-notebook/ Published: 2012-05-12 Categories: ipython, productivity, statistics, tutorial Tags: analysis of variance, ANOVA, bootstrap, confidence interval, ipython, Mann-Whitney-Wilcoxon, MWW, notebook, python, RankSum, research, standard error, statistics, tutorial Randy Olson provides a short demo on how to use IPython Notebook as a research notebook. As promised, here's the IPython Notebook tutorial I mentioned in my introduction to IPython Notebook. Downloading and installing IPython Notebook You can download IPython Notebook with the majority of the other packages you'll need in the Anaconda Python distribution. From there, it's just a matter of running the installer, clicking Next and Accept buttons a bunch of times, and voila! IPython Notebook is installed. Running IPython Notebook For Mac and Linux users, open up your terminal. Windows users need to open up their Command Prompt. Change directories in the terminal (using the cd command) to the working directory where you want to store your IPython Notebook data. To run IPython Notebook, enter the following command: ipython notebook It may take a minute or two to set itself up, but eventually IPython Notebook will open in your default web browser and should look something like this: (NOTE: currently, IPython Notebook only supports Firefox and Chrome.) Creating a new notebook Conveniently, Titus Brown has already posted a quick demo on YouTube. (Start at 2m16s.) Now that we've covered the basics, let's get into how to actually use all this as a research notebook. Using IPython Notebook as a research notebook The great part about the seamless integration of text and code in IPython Notebook is that it's entirely conducive to the "form hypothesis - test hypothesis - evaluate data - form conclusion from data - repeat" process that we all follow (purposely or not) in science. For this example, let's say we're studying an Artificial Life swarm system and the effects of various environmental parameters on the swarm. Here's the example research notebook: [pdf] [ipynb w/ accompanying files] I designed this demo research notebook to be a self-guided tour through the thought process of a researcher as he works on a research project, so hopefully it's helpful to other researchers out there. Statistics in IPython Notebook UPDATE (10/19/2012): Please refer to my other blog post for an up-to-date guide on statistics in Python. For those of you who (understandably) don't want to search through an entire research notebook to figure out how to do statistics in IPython Notebook, here's the cut and dry code. Reading data # Library for reading and parsing csv files import csv # My personal library that contains some useful helper functions import rso_stats # Read and parse data for file "control1.csv" control1 = csv.reader(open('control1.csv', 'rb'), delimiter=',') control1, control1_columns = rso_stats.parse_csv_data(control1) control1 is the dictionary of parsed data control1_columns is the list of column names used to access the data dictionary, sorted in the same order as the csv data file. NOTE: This uses a function from my custom Python library, which parses the data into convenient data dictionaries. The data in the dictionaries can be accessed by: # Access the first column's list of data control1[control1_columns[0]] # Access the fourth column's list of data control1[control1_columns[3]] Standard error of the mean import scipy from scipy import stats mean = scipy.mean(dataset_list) # Compute 2 standard errors of the mean of the values in data_list stderr = 2.0 * stats.sem(dataset_list) Bootstrapped 95% confidence intervals The code below shows you how to compute bootstrapped 95% CIs for the mean. However, this function can bootstrap any range of CIs for any statistical function (mean, mode, standard deviation, etc.). Here's the input parameter description: Input parameters: data = data to get bootstrapped CIs for statfun = function to compute CIs over (usually, mean) alpha = size of CIs (0.05 --> 95% CIs). default = 0.05 n_samples = # of bootstrap populations to construct. default = 10,000 Returns: bootstrapped confidence intervals, formatted for the matplotlib errorbar() function import scipy import rso_stats CIs = rso_stats.ci_errorbar(dataset_list, scipy.mean) NOTE: This uses a couple functions from my custom Python library, since bootstrapping CIs isn't currently supported by SciPy/NumPy. Mann-Whitney-Wilcoxon RankSum test from scipy import stats z_stat, p_val = stats.ranksums(dataset1_list, dataset2_list) Analysis of variance (ANOVA) SciPy's ANOVA function takes two or more dataset lists as its input parameters. from scipy import stats f_val, p_val = stats.f_oneway(dataset1_list, dataset2_list, dataset3_list, ...) Hopefully everyone finds this useful. Get in touch if you have any more ideas on IPython Notebook as a research notebook, or if you'd like to figure out how to do some more statistical tests in Python. --- ## IPython Notebook: Finally, the research notebook I've always been looking for is here! URL: https://www.randalolson.com/2012/05/10/ipython-notebook/ Published: 2012-05-10 Categories: ipython, productivity, statistics Tags: ipython, notebook, python, research Randy Olson discusses the new digital research notebook provided in the IPython library. I attended a Software Carpentry workshop hosted by Titus Brown and Greg Wilson this week and was introduced to, among many other things, a piece of software that I've been looking for ever since I started my graduate program: IPython Notebook. It can easily be installed with the majority of the other packages you'll need in the Anaconda Python distribution. I do the majority of my post-experiment data analysis in Python nowadays, since it's one of the few sanely-designed scripting languages out there with all the functionality I need. What I've been missing is a seamless user interface where I can both take notes about my research and perform my data analysis in the same location. IPython Notebook finally provides that. Ever since I announced my conversion from RTF files to IPython Notebook as my primary means of taking research notes, I've received a lot of flack about how Python doesn't support advanced statistical tests, such as bootstrapping confidence intervals, Mann-Whitney Wilcoxon RankSum tests, and ANOVA tests. After a day of searching with my lab mates, I finally turned up all the libraries I need: Bootstrapped confidence intervals: https://pypi.python.org/pypi/scikits.bootstrap MWW RankSum test: http://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.ranksums.html ANOVA: http://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.f_oneway.html If you can't find a Python library for a statistical test you need, post here and we'll try to find it. The IPython notebook has plenty of uses beyond a research notebook, too. For example, Titus Brown recently posted the IPython notebook that he used to generate all of the graphs in one of his recent papers. Imagine the implications for science if scientists actually start showing the code they used to generate their graphs! (No more hiding that outlier point on the side of the graph...) I'll be putting up some tutorials and examples of how to use IPython Notebook for exploratory statistical data analysis soon, so stay posted! ---