Releasing open weights that can approach frontier level intelligence (irrespective of number of tokens burned) is just a way of telling the world that anyone, even China, can serve frontier level inference if they have the chips and warm shells to do so.
What is stopping China from gaining a majority market share, then, in terms of serving inference?
AI Sovereignty -- yes
Cybersecurity concerns -- yes
Latency -- no, unlike previous emerging IT workload types , inference does not have strong latency requirements. eg 1s of additional network latency doesn't matter to a 15 min, 10-turn agent session.
Cost -- ultimately this comes down to a nations ability to plug chips into warm shells. which forks into geopolitical / trade on the chips side and energy scalability and modularity on the warm-shell side. Even if you call geopolitical / trade a toss-up, China has the US beat HANDILY on the energy front, yearly they are deploying 10x power to their grid relative to the US, which is shooting itself in the foot at every possible moment.
IMHO chip tech will travel across borders, absent a breakthrough in analog inference, energy scalability will ultimately dominate.
Seems like a common-sense approach. I appreciate the emphasis on understanding, humans will eventually be held accountable, blaming Claude for an outage is not going to get Claude fired.
Is this LBO a ridiculous enough of a proposition that you are going to email your senator and or representatives to complain about LBO loopholes in corporate finance law?
Add on the compounding effect that "QA" or "test" in someone's job description was viewed as a synonym for "less-highly compensated" over the past few decades, and you have an entire generation of mid career devs with poorly adapted instincts regarding what is valuable in the process of shipping working product.
What is stopping China from gaining a majority market share, then, in terms of serving inference?
AI Sovereignty -- yes
Cybersecurity concerns -- yes
Latency -- no, unlike previous emerging IT workload types , inference does not have strong latency requirements. eg 1s of additional network latency doesn't matter to a 15 min, 10-turn agent session.
Cost -- ultimately this comes down to a nations ability to plug chips into warm shells. which forks into geopolitical / trade on the chips side and energy scalability and modularity on the warm-shell side. Even if you call geopolitical / trade a toss-up, China has the US beat HANDILY on the energy front, yearly they are deploying 10x power to their grid relative to the US, which is shooting itself in the foot at every possible moment.
IMHO chip tech will travel across borders, absent a breakthrough in analog inference, energy scalability will ultimately dominate.
reply