Open-Source AI Statistics in US 2026 | Models, Users & Key Facts

Open-Source AI in America 2026

Open-source AI has moved from a developer niche to a genuine pillar of the American technology stack in 2026, even as the country’s biggest labs continue to dominate headlines with closed, proprietary systems. According to the Linux Foundation, 89% of U.S. organizations that have adopted AI now use open-source models or tools somewhere in their infrastructure, and open-source components make up roughly 41% of the average company’s AI stack. At the same time, Stanford’s 2026 AI Index shows the United States still leads the world in notable model releases, producing 59 in 2025 compared with China’s 35 — but only a small handful of those American releases were fully open-sourced, underscoring a split between America’s open developer ecosystem and its closed frontier labs.

That split defines the current moment for open-source AI in the US. On one hand, Meta’s Llama family has passed 1.2 billion cumulative downloads, and Hugging Face now hosts more than 2.8 million model repositories, evidence of an enormous and still-growing open developer community. On the other hand, enterprise buying data tells a more cautious story: according to Menlo Ventures’ third annual survey of U.S. enterprise AI decision-makers, open-source models’ share of production LLM API usage actually fell between 2024 and 2025, even as overall enterprise AI spending nearly tripled. Understanding 2026’s open-source AI landscape means holding both of these trends at once — explosive grassroots and developer adoption alongside more measured enterprise deployment at the largest companies, a tension that shapes nearly every statistic covered below.

Interesting Facts About Open-Source AI in the US 2026

Fact Detail
US Notable AI Models Released, 2025 59, versus China’s 35
Fully Open-Sourced Models Among 95 “Important” 2025 Releases 4
Foundation Model Transparency Index Average Fell to 40 (from 58 the prior year)
Meta Llama Cumulative Downloads 1.2 billion+
Hugging Face Model Repositories Listed 2,835,314
Open-Source Share of Enterprise LLM API Usage, 2025 11%, down from 19% in 2024
US Enterprises Using Some Open Source in AI Stack 89% of AI adopters
GitHub Open-Source AI Projects Tracked 5.6 million
Performance Gap, Top Closed vs. Top Open Model (March 2026) 3.3–3.4% on Arena rankings
OpenRouter Routed Tokens, Nov 2024–Nov 2025 100+ trillion

Source: Stanford HAI 2026 AI Index; Linux Foundation Research; Meta; Hugging Face; Menlo Ventures; GitHub Octoverse; OpenRouter

The gap between 59 US model releases and just 4 fully open-sourced releases among 2025’s most important models captures the central tension in American AI right now: the country that produces the most frontier models is, at the very same time, disclosing less about how those models work than it did a year earlier. Stanford’s own Foundation Model Transparency Index — which tracks whether labs publish training code, dataset details, and parameter counts — fell from an average of 58 to 40 in a single year, driven largely by leading labs pulling back on disclosure even as their commercial and strategic importance grew.

Meanwhile, the developer-facing numbers point in the opposite direction entirely. Meta’s Llama crossing 1.2 billion downloads, combined with Hugging Face’s 2.8 million-plus repositories and GitHub’s 5.6 million tracked open-source AI projects, shows a developer community that has never had more raw material to work with. The seeming contradiction between a declining enterprise open-source API share and surging download and repository counts is best explained by where each number is measured: downloads capture curiosity, experimentation, and small-scale self-hosting, while enterprise API share captures which models power a company’s largest, highest-stakes production workloads — and at that top tier, many US enterprises are still choosing closed models from OpenAI, Anthropic, and Google.

Open-Source AI Model Landscape Statistics in the US 2026

Metric 2025–2026 Data
US Notable Model Releases, 2025 59
China Notable Model Releases, 2025 35
Share of Notable Models From Industry (vs. Academia) 90%+
Fully Open-Sourced Models Among 95 Important 2025 Releases 4
Foundation Model Transparency Index, 2025 Average 40 (down from 58 in 2024)
Top Closed vs. Top Open Model Gap (March 2026, Arena Elo) 3.3–3.4%
Narrowest Recorded Open-Closed Gap (August 2024) 0.5%

Source: Stanford HAI 2026 AI Index; Epoch AI

The United States’ lead in raw model output — 59 notable releases against China’s 35 — is real, but it obscures how concentrated and closed that output has become. Stanford’s 2026 AI Index found that industry now produces more than 90% of all notable models, a sharp shift away from the academic labs that once drove much of American AI research, and within that industry-dominated pipeline, openness has become the exception rather than the rule: only 4 of the 95 models the report classifies as “important” in 2025 were released as fully open. That is a striking reversal from the field’s earlier years, when open releases from labs like Meta and Mistral were treated as a competitive necessity rather than a rare exception.

The performance gap between the best open and best closed models offers a useful reality check against any narrative of open-source falling behind. At 3.3 to 3.4% as of March 2026, the gap between top closed models and the strongest open-weight competitors remains narrow by historical standards, even after widening slightly from the 0.5% low point recorded in August 2024. In practice, this means most everyday tasks show little measurable difference between a leading open model and a leading closed one — the gap tends to matter most at the very frontier of reasoning and coding benchmarks, which is exactly where the largest US labs have concentrated their disclosure cutbacks. Stanford’s researchers have noted that reported parameter counts have hovered near a trillion for roughly three years running, even as independently estimated training compute for the newest frontier systems keeps climbing, a further sign that public disclosure has decoupled from the actual scale of what labs are building.

Open-Source AI Enterprise Adoption Statistics in the US 2026

Metric Data
US Organizations That Have Adopted AI Tools 94%
AI Adopters Using Some Open Source in Their Stack 89%
Average Share of AI Infrastructure That Is Open Source 41%
Organizations Expecting Open-Source Use to Increase 76%
Organizations Citing Open Source as Cheaper to Deploy ~67% (two-thirds)
Organizations Choosing Open Source Specifically for Cost Savings ~48% (nearly half)
Open-Source Share of Enterprise LLM API Usage, 2024 → 2025 19% → 11%
Enterprise Generative AI Spending, 2024 → 2025 $11.5B → $37B

Source: Linux Foundation Research; Menlo Ventures (2025 Enterprise AI Report, 495 US decision-makers surveyed)

The Linux Foundation’s research paints a picture of open source as deeply embedded infrastructure: 94% of surveyed US organizations have adopted some form of AI, and 89% of those adopters rely on open-source components somewhere in that stack, whether for embeddings, fine-tuning frameworks, or full model deployment. Cost remains the dominant driver — roughly two-thirds of organizations say open-source AI is cheaper to run than proprietary alternatives, and nearly half specifically cite cost savings as their reason for choosing it, a rational response given that closed-model API pricing can run several times higher per token than comparable open alternatives.

Yet Menlo Ventures’ enterprise-specific data, drawn from 495 US enterprise AI decision-makers surveyed in late 2025, tells a more complicated story about where that open-source infrastructure actually gets deployed. Open-source models’ share of production LLM API usage fell from 19% to 11% between the 2024 and 2025 surveys — even as total enterprise generative AI spending more than tripled, from $11.5 billion to $37 billion. Read together, these figures suggest enterprises are pouring far more money into AI overall while concentrating an increasing share of their highest-value, highest-visibility workloads on closed frontier models, reserving open-source deployments for smaller, cost-sensitive, or infrastructure-level tasks rather than flagship production systems. The apparent contradiction between rising adoption surveys and falling API-usage share is largely a measurement question: adoption surveys count any organization using open source anywhere, even for a single minor internal tool, while API-usage share reflects the proportion of actual production traffic, weighted by application scale — a distinction that matters enormously when interpreting any single open-source statistic in isolation.

Open-Source AI Model Downloads and Repository Statistics in the US 2026

Metric Data
Hugging Face Model Repositories 2,835,314
Top 200 Models’ Share of All Hugging Face Downloads 49.6%
Meta Llama Cumulative Downloads (as of early 2026) 1.2 billion+
Llama Downloads, December 2024 650 million
GitHub-Tracked Open-Source AI Projects 5.6 million
Of Those, Projects With 10+ GitHub Stars 206,880
Increase in GitHub Generative AI Projects, 2024 98%
Most-Downloaded Single Model (all-MiniLM-L6-v2, embeddings) 255 million downloads

Source: Hugging Face; Meta; GitHub Octoverse; Stanford HAI 2026 AI Index

Meta’s Llama remains the flagship story of American open-source AI distribution — the jump from 650 million downloads in December 2024 to 1.2 billion by early 2026 represents nearly a doubling in little more than a year, and Meta has said the milestone reflects thousands of developers contributing tens of thousands of derivative models, which are themselves downloaded hundreds of thousands of times each month. That said, raw download counts concentrate heavily at the top: Hugging Face’s own data shows the top 200 models capture nearly half (49.6%) of all platform downloads, meaning the long tail of millions of smaller repositories splits the remaining share extremely thinly.

GitHub’s numbers add another dimension to the picture — its 5.6 million tracked open-source AI projects sound enormous until compared with the 206,880 that have attracted 10 or more stars, a rough proxy for genuine community traction rather than abandoned experiments or auto-generated forks. The 98% surge in generative AI projects that GitHub recorded in 2024 suggests this gap between total projects and meaningfully-adopted ones will likely persist or widen as the barrier to publishing a repository keeps falling faster than the barrier to building something developers actually rely on. It’s also worth separating downloads from genuine production usage: a high download count says a model is widely fetched, not necessarily widely deployed, since continuous-integration jobs, mirror syncs, and container rebuilds all register as separate downloads even when a single team or pipeline is responsible for repeated pulls. For a wider view of how this open-source download and repository activity fits into the country’s overall AI adoption picture, the Artificial Intelligence Statistics in US report tracks usage, adoption, and investment trends across both open and closed AI in the US.

Open-Source AI Token Usage and Platform Statistics in the US 2026

Metric Data
OpenRouter Routed Tokens, Nov 2024–Nov 2025 100+ trillion
OpenRouter Annualized Run Rate, May 2026 ~1.5 quadrillion tokens
OpenRouter Developer Count Growth (roughly 1 year) 2.5 million → 8 million+
Open-Weight Share of OpenRouter Token Volume, Late 2025 ~33%
Chinese Open-Weight Share of OpenRouter Traffic, April 2026 ~45% of weekly tokens
DeepSeek Tokens Processed, Nov 2024–Nov 2025 14.37 trillion
Qwen Derivative Models on Hugging Face (March 2026) 113,000+

Source: OpenRouter State of AI 2025; Menlo Ventures; Hugging Face

OpenRouter’s growth trajectory is one of the clearest signals that open-weight model usage is scaling rapidly even if it isn’t dominating enterprise API budgets. The platform’s developer base grew from 2.5 million to more than 8 million in roughly a year, while its annualized token run rate climbed to an estimated 1.5 quadrillion tokens by May 2026 — a scale that would have been difficult to imagine just two years earlier. Within that traffic, open-weight models accounted for roughly a third of all routed tokens by late 2025, a meaningfully larger share than the enterprise API figures from Menlo Ventures alone would suggest, since OpenRouter’s user base skews toward developers and smaller companies experimenting more freely with open alternatives.

What stands out most inside that open-weight traffic is how much of it now flows to Chinese labs rather than American ones. DeepSeek alone processed 14.37 trillion tokens on OpenRouter across the year studied, and by April 2026, Chinese open-weight models — led by DeepSeek and Alibaba’s Qwen family — accounted for roughly 45% of weekly OpenRouter traffic, up from under 2% in late 2024. Halfway through this shift, Qwen also became the dominant ecosystem for derivative development on Hugging Face, with more than 113,000 direct derivative models built on top of it by March 2026 — a figure that now exceeds Llama’s derivative count and signals a real change in which open-source family US developers are building around.

Open-Source AI Investment and Funding Statistics in the US 2026

Metric Data
Hugging Face Valuation $4.5 billion
Hugging Face Total Funding Raised $400 million
Mistral AI Valuation, September 2025 €11.7B (~$13.5B)
Mistral AI Valuation in Talks, Mid-2026 ~€20B (~$23B)
Total AI Startup Funding, 2025 ~$150 billion
Share of Global VC Captured by AI Startups, 2025 40%+
Foundation Model Companies’ Share of 2025 AI Funding $80 billion

Source: Contrary Research; TechCrunch; Bloomberg; Wellows AI Startup Rankings

Hugging Face, the New York- and Paris-based hub that hosts the majority of the open-source AI ecosystem’s models and datasets, has raised a comparatively modest $400 million against a $4.5 billion valuation — a reflection of its role as critical infrastructure rather than a model developer competing directly at the frontier. Mistral AI, by contrast, has seen its valuation move dramatically even though it is headquartered in Paris rather than the US: from roughly $13.5 billion in September 2025 to reported talks near $23 billion by mid-2026, driven substantially by American investors including Andreessen Horowitz, General Catalyst, and Nvidia, illustrating how deeply US capital remains embedded in open-model development even when the labs themselves sit outside American borders.

Set against the broader AI funding environment, these open-source-focused rounds are still a small fraction of the total. AI startups collectively raised roughly $150 billion in 2025, representing more than 40% of all global venture capital deployed that year, with foundation model companies alone capturing $80 billion of that total. For a fuller picture of where that capital is concentrating across the AI sector overall — not just among open-source specialists — the AI Investment Statistics in US report tracks the broader funding trends behind these individual company numbers.

Small Language Model and Infrastructure Statistics in the US 2026

Metric Data
Global Small Language Model Market Size, 2026 $10.65–$11.1 billion
Small Language Model Market Size, 2025 $0.93 billion
Most-Downloaded Hugging Face Model (embeddings) 255 million downloads
Leading Text-Generation Model Downloads (Qwen3-0.6B) 28.3 million
Closed-Model Per-Token Cost vs. Open Models ~6x higher
Total Cost of Ownership Reduction From Shifting to Open Models ~35%

Source: MarketsandMarkets; Hugging Face; industry cost-comparison analyses

The explosive growth projected for the small language model market — from under $1 billion in 2025 to more than $10 billion in 2026 by several research firms’ estimates — reflects a growing recognition among US developers that not every AI task requires a massive, expensive frontier model. Small, efficient open-weight models are increasingly favored for narrowly scoped tasks like classification, embeddings, and lightweight chat features, which explains why the single most-downloaded model on Hugging Face isn’t a chatbot at all, but a compact 22.7-million-parameter embedding model that has been pulled 255 million times.

Cost economics continue to be the clearest argument in open source’s favor at this smaller scale: closed frontier models can cost roughly six times more per token than comparable open alternatives, and companies that shift applicable workloads to open models report total cost of ownership reductions of around 35%. Running these smaller, self-hosted models still requires real infrastructure, however, and the computing capacity needed to train and serve them at scale continues to expand rapidly across the country. Readers interested in where that physical infrastructure is being built can find further detail in the Data Center Statistics in US report, which tracks the buildout supporting both the open and closed sides of America’s AI industry.

Disclaimer: This research report is compiled from publicly available sources. While reasonable efforts have been made to ensure accuracy, no representation or warranty, express or implied, is given as to the completeness or reliability of the information. We accept no liability for any errors, omissions, losses, or damages of any kind arising from the use of this report.