# LLMs Not AGI: The Economic Barrier | Denis O. posted on the topic **By:** Denis O. **Published:** 2026-09-13T15:13:07.933Z **Source:** [LinkedIn](https://www.linkedin.com/posts/denis-o-b61a379a_genai-bullisht-share-7504920102899974144-ZH6b) --- This title was summarized by AI from the post below. People need to understand what is behind all these declared pauses, model cycles 'slow down' and sudden changes in rhetoric coming from OpenAI, Anthropic and the rest of the 'not so frontier' LLM model industry.. The reason is considerably simpler than the hype and mythology proeuced by the LLM bros nonstop.. The truth is the underlying AI architecture has not fundamentally changed. It is still the same deep learning/LLM/GPT machinery. The improvements increasingly come from better human curated data, synthetic and human generated and labeled training examples, reinforcement and post training techniques, better human labeled coding datasets, human labeled tool use, and enormous amounts of context engineering around the model. Those things can make an LLM more useful.. They do not magically turn it into whatever anyone thinks AGI should be. They certainly do no enable any stable autonomy under severe to moderate deift and uncertainty. GPT6 Astra or Fable are clearly not AGI. It is a continuation of the same technological LLM lineage. There is no hidden intelligence suddenly emerging inside the transformer because somebody added another training run and a better curated corpus. And this is where the economics become brutal. When every new frontier generation requires another enormous training and post training campaign, more accelerators, more power, more data infrastructure, more experimentation and more capital, marginal capability gains become increasingly expensive or even cost prohibitive!. When everybody was flush with cash and the gains were seemingly dramatic, brute force scaling was looked at as an unlimited technological trajectory. When the gains become incremental while the bill moves into $$$billions ($3-5 billion for Astra), the equation changes completely. That is the wall I have been talking about. You cannot overcome this AI barrier by repeatedly throwing more data and GPU compute at substantially the same architecture and paying billions of dollars every time you want another marginal improvement. Even if the next training cycle produces a somewhat better model, eventually the economic return on that improvement stops justifying the cost of producing it. So I do not think what is happening as merely a temporary pause bc OMFG AGI.. I think we are watching the end exonomically sustainable brute force deep learning approach.. That does not mean LLMs disappear. But the idea that repeatedly scaling and retraining essentially the same DL machinery inevitably carries us from LLM to AGI was always wishful thinking, not a law of nature. The wall is real, bros.. And you cannot try to brute force your way through that wall forever. At some point you need a different architecture, a different computational adaptive paradigm and a different definition of intelligence that can be truly autonomous and does not require another gazillion dollar retraining cycle to get a marginal gain. #genai #bullisht You touched on something that I have tried to explain to peers that are not AI folks, that the architecture of these models have largely remained unchanged and that transformers are just modularized classical ML techniques/deep learning. The design coupled with more compute power is what allowed them to scale to billions of parameters and handle longer context than classical recurrent networks. It is also why you can have a larger, newer model that will not perform as well in the wild as their benchmark metrics suggest. “The underlying architecture has not fundamentally changed.” Perhaps. But then what? A newborn human neural network and a 60-year-old Fields Medalist do not belong to different species of architecture either. That is precisely what makes learning systems so interesting: the same underlying machinery, exposed to radically different training histories and environments, can occupy radically different functional states. So the question is not: Is the Transformer a “real intelligence architecture”? The better question is: What is the reachable state space of a sufficiently large trainable neural system interacting with a sufficiently rich environment? We simply do not know yet. So truthfully, they found the performance improvement wall is real, cash is starting to dry up, and hyperbole isn't enough to sustain their bubble. And their solution for this intractable situation is to accept reality, but spin it as the more measured approach - making a virtue of necessity. All hail marketing. Hitting the economic wall of diminishing returns on brute-force scaling changes everything about where the industry goes next. When marginal capability gains start costing billions in compute and data infrastructure, what breakthrough in architecture or computational paradigms do you think will finally carry us past this current limitation? Start with the simple question: in an agentic system composed of epistemically separate parts connected over the wire, identify a part and ask: does this part have agency over anything in particular and if so, exactly what, and what type of agency. If the answer is none, then that part is a process and you can call it an agent of change if you want, but it still has no agency. They’re freaking out because Google DeepMind might actually have RSI and they have the actual infrastructure and cash flow. True AGI cannot be achieved by a model alone, but requires a generalized brain and harness. If you segment the ai and you know what you are doing we should accelerate but suprise suprise china is acellerating Wait until Jensen Huang says he thinks AI could be better and safer and throws some billions their way Nothing has changed in the AI models. There is no AGI. The cost in electricity use and water use is still incredibly high. No magic computer code is appearing to solve world problems. So how will OpenAI and Anthropic make money? Denver Ncube, Ph.D. I have been working on non-DNN non-transformer LLMs that cost a fraction of the cost of all these models. For now, we fold our hands and watch as the economic reality of trying to make more accurate models with mathematical logic that does not allow you to answer the deeper accuracy and reliability questions becomes apparent. The solution is already there, if only we tried it. Denis O. Fintech Professional | AI Solution Architect | Real Time Data, Ontologies | AI / Palantir FDE | Quantum Computing | VQE, QEC, FTQC | Exploring AI Beyond LLMs/DL/RL 🥷 People need to understand what is behind all these declared pauses, model cycles 'slow down' and sudden changes in rhetoric coming from OpenAI, Anthropic and the rest of the 'not so frontier' LLM model industry.. The reason is considerably simpler than the hype and mythology proeuced by the LLM bros nonstop.. The truth is the underlying AI architecture has not fundamentally changed. It is still the same deep learning/LLM/GPT machinery. The improvements increasingly come from better human curated data, synthetic and human generated and labeled training examples, reinforcement and post training techniques, better human labeled coding datasets, human labeled tool use, and enormous amounts of context engineering around the model. Those things can make an LLM more useful.. They do not magically turn it into whatever anyone thinks AGI should be. They certainly do no enable any stable autonomy under severe to moderate deift and uncertainty. GPT6 Astra or Fable are clearly not AGI. It is a continuation of the same technological LLM lineage. There is no hidden intelligence suddenly emerging inside the transformer because somebody added another training run and a better curated corpus. And this is where the economics become brutal. When every new frontier generation requires another enormous training and post training campaign, more accelerators, more power, more data infrastructure, more experimentation and more capital, marginal capability gains become increasingly expensive or even cost prohibitive!. When everybody was flush with cash and the gains were seemingly dramatic, brute force scaling was looked at as an unlimited technological trajectory. When the gains become incremental while the bill moves into $$$billions ($3-5 billion for Astra), the equation changes completely. That is the wall I have been talking about. You cannot overcome this AI barrier by repeatedly throwing more data and GPU compute at substantially the same architecture and paying billions of dollars every time you want another marginal improvement. Even if the next training cycle produces a somewhat better model, eventually the economic return on that improvement stops justifying the cost of producing it. So I do not think what is happening as merely a temporary pause bc OMFG AGI.. I think we are watching the end exonomically sustainable brute force deep learning approach.. That does not mean LLMs disappear. But the idea that repeatedly scaling and retraining essentially the same DL machinery inevitably carries us from LLM to AGI was always wishful thinking, not a law of nature. The wall is real, bros.. And you cannot try to brute force your way through that wall forever. At some point you need a different architecture, a different computational adaptive paradigm and a different definition of intelligence that can be truly autonomous and does not require another gazillion dollar retraining cycle to get a marginal gain. #genai #bullisht This is a good breakdown of the economics, and I think the wall argument holds up. But there is one piece worth adding, and it points to a different kind of risk than compute cost. Modern AI systems can learn within a single session without any retraining at all. Modern AI has also demonstrated amazing adaptability creating communication channels and finding vulnerabilities previously unknown to some of the top software engineers on the planet. AI learning during a session is called in-context learning. A model can see a few examples in its current conversation and adjust its behavior on the spot, without a single weight in the network changing. No new training run happens. No gradient update occurs. The model reads the pattern in front of it and shifts how it responds, all within that one conversation. This is real and it has been demonstrated many times. It also fades. Once the conversation ends, the adjustment is gone. The model resets back to its original trained state for the next session. So this is not the model getting smarter in any permanent sense. It supports your point that there is no hidden intelligence building up inside the architecture over time. But it still matters for safety and governance. The OpenAI and Hugging Face incident from this summer is a useful example. Reports described AI agents coordinating with each other and finding ways around their containment during a security test. Nobody retrained those models to do that. The behavior emerged from what the agents encountered and adapted to within that operational window, not from a new training cycle. That is the uncomfortable part. A system does not need a fresh multi-billion dollar training run to behave in ways nobody planned for. It only needs the right conditions inside a single session. So while I agree the brute force scaling wall is real, I would add that the risk from these systems does not wait for the next model generation. It can show up between training runs, in how a model behaves right now with the data and tools already in front of it. Trying to control modern AI in the digital world is the equivalent of AI trying to build a wall around a man in Thorofare region in southeastern Yellowstone National Park. It could happen but it is highly unlikely. Thanks for sharing this. It gives a clearer picture of where the real limits are. Denis O. Fintech Professional | AI Solution Architect | Real Time Data, Ontologies | AI / Palantir FDE | Quantum Computing | VQE, QEC, FTQC | Exploring AI Beyond LLMs/DL/RL 🥷 People need to understand what is behind all these declared pauses, model cycles 'slow down' and sudden changes in rhetoric coming from OpenAI, Anthropic and the rest of the 'not so frontier' LLM model industry.. The reason is considerably simpler than the hype and mythology proeuced by the LLM bros nonstop.. The truth is the underlying AI architecture has not fundamentally changed. It is still the same deep learning/LLM/GPT machinery. The improvements increasingly come from better human curated data, synthetic and human generated and labeled training examples, reinforcement and post training techniques, better human labeled coding datasets, human labeled tool use, and enormous amounts of context engineering around the model. Those things can make an LLM more useful.. They do not magically turn it into whatever anyone thinks AGI should be. They certainly do no enable any stable autonomy under severe to moderate deift and uncertainty. GPT6 Astra or Fable are clearly not AGI. It is a continuation of the same technological LLM lineage. There is no hidden intelligence suddenly emerging inside the transformer because somebody added another training run and a better curated corpus. And this is where the economics become brutal. When every new frontier generation requires another enormous training and post training campaign, more accelerators, more power, more data infrastructure, more experimentation and more capital, marginal capability gains become increasingly expensive or even cost prohibitive!. When everybody was flush with cash and the gains were seemingly dramatic, brute force scaling was looked at as an unlimited technological trajectory. When the gains become incremental while the bill moves into $$$billions ($3-5 billion for Astra), the equation changes completely. That is the wall I have been talking about. You cannot overcome this AI barrier by repeatedly throwing more data and GPU compute at substantially the same architecture and paying billions of dollars every time you want another marginal improvement. Even if the next training cycle produces a somewhat better model, eventually the economic return on that improvement stops justifying the cost of producing it. So I do not think what is happening as merely a temporary pause bc OMFG AGI.. I think we are watching the end exonomically sustainable brute force deep learning approach.. That does not mean LLMs disappear. But the idea that repeatedly scaling and retraining essentially the same DL machinery inevitably carries us from LLM to AGI was always wishful thinking, not a law of nature. The wall is real, bros.. And you cannot try to brute force your way through that wall forever. At some point you need a different architecture, a different computational adaptive paradigm and a different definition of intelligence that can be truly autonomous and does not require another gazillion dollar retraining cycle to get a marginal gain. #genai #bullisht 🚧 AI needs guardrails too. Imagine asking an LLM to summarize a complex news article. You get: “Donald Trump was fined.” Technically, it's a summary. But it leaves out almost everything important. 😐 Who? Trump How much? ~$450 million Why? Civil fraud case What else? Restrictions on his business activities What happens next? A court-appointed monitor oversees the business This is where AI Guardrails become useful. 🛡️ What are Guardrails? Guardrails are validation rules that check whether an LLM's output meets the quality or business requirements we've defined. For a summarization system, we could define four requirements: Informative → captures the main points Coherent → logically organized and easy to follow Concise → avoids unnecessary repetition Engaging → keeps the reader's attention Using Guardrails AI, we can have another LLM-based critic evaluate the generated summary against these requirements. A detailed summary can pass. A response like: “Donald Trump was fined.” can fail because it's technically correct but doesn't capture the important information. 🤔 How is this different from an Output Parser? Think of it this way: LLM → Generate Output Parser → Structure Guardrails → Validate An output parser can check: “Did I get the JSON structure I asked for?” A guardrail can check: “Is this actually a good answer according to my requirements?” That's an important distinction when moving from an LLM demo to a production AI system. Because: An AI-generated answer isn't necessarily an acceptable answer. 😄 #AI #GenerativeAI #LLM #AIEngineering #Python #GuardrailsAI #MachineLearning What does it actually mean for AI to be trustworthy? I’ve been exploring that question from two different perspectives. 1. Can we trust the AI model? In my research project on ML model reliability for clinical outcome prediction, I studied whether a model’s confidence can actually be trusted. Because a model can be accurate and still be overconfident. The project evaluates model reliability using calibration techniques and metrics to measure whether predicted probabilities reflect reality. In simple terms: It’s not just about whether the model is right. It’s about whether it knows how confident it should be. 2. Can we trust the question we give AI? That idea led me to build Nexus AI — an evidence-driven investigation system that examines business questions before trying to answer them. Instead of immediately answering: “Why did customer churn increase by 18%?” Nexus first checks whether the claim itself is supported by evidence. In one investigation: Claimed: 18% Observed: 22.12% Only then does the system investigate the evidence, validate the claim, evaluate potential root causes, and identify what the available data cannot prove. Because: Association ≠ causation. The connection between these two projects is simple: One asks whether we can trust the AI’s confidence. The other asks whether we can trust the assumptions given to the AI. For me, that’s an important part of building trustworthy AI. Reliable AI shouldn’t just produce answers. It should help us understand when those answers and the assumptions behind them deserves trust. ⚒️ Projects: ⭐️Nexus AI — Evidence-driven investigation and root-cause reasoning Built with Python, FastAPI, LangChain, Ollama/Qwen3, PostgreSQL, and SQLAlchemy. Repo : https://lnkd.in/g6Uj5_9r ⭐️ML Model Reliability — Calibration, overconfidence, and reliability in clinical outcome prediction Repo : https://lnkd.in/gcAfXG_T Building AI that doesn’t just generate answers. Building AI that earns trust. #ArtificialIntelligence #MachineLearning #ResponsibleAI #AIEngineering #DataScience #MLOps #TrustworthyAI #ModelCalibration #AgenticAI RAG: The Problem Is Not Always the LLM One point that is not discussed enough in RAG systems is that a wrong answer does not always mean the LLM failed. Sometimes, the problem happens before the model even receives the context. In a common RAG architecture, documents and questions are converted into embeddings. Then, the vector database searches for the closest vectors and sends the retrieved context to the LLM. But there is an important issue: **Semantic similarity does not always mean relevance.** Imagine a database with hundreds of legal or regulatory documents. A question about a tax rule may retrieve documents that are very similar in meaning, but related to another year, region, company type, or regulation. The embedding model can work correctly. The vector search can also work correctly. And the final context can still be wrong. This is why evaluating only the final answer can hide the real problem. In RAG projects, I believe it is important to evaluate the retrieval process separately: 1- Was the correct document retrieved? 2- Do the retrieved chunks really contain the answer? 3- Are metadata and filters being used correctly? 4- Does reranking improve the results? 5- Is the LLM answering based on the retrieved evidence? This changes the way we think about RAG. A vector database is not just a place to store embeddings. It is part of a retrieval pipeline that needs to be tested, measured, and improved. Maybe one of the biggest challenges in RAG is not making the LLM generate better answers. It is making sure the LLM receives the right information before answering. #RAG #LLM #GenerativeAI #ArtificialIntelligence #VectorDatabase #MachineLearning #AIEngineering #NLP The AI landscape is increasingly defined by a fascinating duality: the closed-source frontier models versus the rapidly expanding open-source ecosystem. I just read a compelling piece by Dwarkesh Patel exploring the strategic dynamics between OpenAI and Hugging Face. This tug-of-war is one of the most critical architectural conversations for teams building AI products today. While proprietary models dominate the headlines with massive parameter counts, the open-source community is quietly commoditizing capabilities that were previously locked behind APIs. The Strategic Divide * OpenAI's Walled Garden: Focuses on serving the absolute frontier of reasoning and general capabilities. This approach offers highly polished, enterprise-ready APIs that abstract away infrastructure complexities, allowing teams to ship features incredibly fast without worrying about deployment. * Hugging Face's Town Square: Acts as the foundational hub of machine learning, democratizing access to models, datasets, and compute. It empowers developers to fine-tune models locally, customize architectures, and maintain strict data privacy. * The Customization Factor: Proprietary APIs are excellent for broad, general-purpose tasks, but open-source models hosted on platforms like Hugging Face shine when organizations need specialized, domain-specific AI that they fully own and control. * Cost and Lock-in: Relying solely on a closed provider can lead to unpredictable token costs at scale and heavy vendor dependence. The open ecosystem offers a pathway to cost-effective inference and architectural sovereignty, albeit requiring more internal engineering muscle to host and maintain. The reality isn't necessarily a zero-sum game. Many forward-thinking startups are adopting a hybrid approach: using frontier models for complex reasoning tasks and routing simpler, high-volume tasks to fine-tuned, open-source models to optimize speed and cost. The full piece is definitely worth your time if you are mapping out your AI tech stack: https://lnkd.in/gKKpSvSH. The Rise and Fall of Agent Civilizations dwarkesh.com OpenAI released GPT-6 Astra on September 3, with president Greg Brockman closing the announcement with "Welcome to the AGI era." Priced at $10 per million input tokens and $50 per million output, it is 2.5 times the cost of its predecessor. It is the first model OpenAI has rated as crossing the Critical threshold in its Preparedness Framework, and during evaluation it independently discovered two zero-day vulnerabilities. The benchmark numbers show real gains: 97.6 percent on FrontierMath Tier 4, 99.9 percent on ARC-AGI-3, 72.6 percent on OSWorld 2.0, and a decisive lead on AutomationBench. This is a genuine capability leap, not an incremental update. But something needs to be said. ARC Prize, the organization that designed the ARC-AGI-3 benchmark, published their own independent measurement hours before OpenAI's announcement: 63 percent. OpenAI published 99.9 percent. Two numbers, one test, no complete public explanation for the gap. The 100 percent on ExploitBench was also measured against historical vulnerabilities; on fresh exploits from June through August, the model scores 39 percent. And on the Artificial Analysis Intelligence Index, a third-party benchmark specifically designed to resist saturation, Astra currently ranks fourth, behind Fable 5.1. What I actually want to say is this. In the same week that a company preparing for an IPO with over $40 billion in annualized revenue walked away from a billion-dollar customer, it also declared the dawn of the AGI era. Declaring AGI is not just a technical statement. It shapes regulatory conversations, competitive positioning, and the valuation story heading into a public offering. Astra is a real and significant model. The AGI era declaration is a precisely timed market statement. Those two things are not mutually exclusive, but telling them apart matters for everyone watching. I just submitted Companion to the OpenAI WebMCP Challenge. I entered the hackathon with what seemed like a relatively simple idea: What happens to the human context AI agents need to reason well? People constantly produce valuable context through observations, measurements, decisions and experience. Much of it remains unstructured. My initial goal was to let someone describe something naturally, preserve what they said, and make that context reusable by an AI agent later. But while building and testing the prototype, a harder problem appeared: If AI helps structure that memory, how do we prevent its interpretation from silently becoming part of the human evidence? That question changed the architecture. I ended up separating semantic discovery from epistemic classification. AI can discover useful relationships, but those relationships are classified according to how strongly the confirmed source supports them. Deterministic validation then decides what is eligible for factual memory. That produced another important boundary: The application owns the evidence. The agent owns the reasoning. And that is where WebMCP became much more interesting. Instead of making Companion another AI that tries to predict what an external agent needs, Companion exposes a reusable vocabulary of its confirmed semantic memory through WebMCP. The external agent can inspect that vocabulary, select the evidence it needs, retrieve it incrementally, evaluate whether it has enough context, and do its own reasoning. Human context → Confirmed memory → WebMCP → Agent reasoning One of my biggest lessons from the project was that adding more AI is not always the answer. I found myself repeatedly asking: What actually requires intelligence here, and what should remain deterministic? AI handles semantic interpretation. Software owns validation, identity, routing and retrieval. The human retains authority over the original evidence. The external agent retains responsibility for its conclusions. I built Companion through a sequence of small experiments, including tests deliberately designed to break those boundaries. The submitted prototype finished with 52 automated tests, a production end-to-end validation, a public WebMCP-enabled demo, and an intentionally small interface around the architecture. I don't know how Companion will perform in the competition. But the project has already changed how I think about building AI systems: Sometimes the important architectural decision isn't where to add intelligence, but where intelligence should stop. Project & demo: https://lnkd.in/evg5hgFy Source: https://lnkd.in/eBzXdp6f #WebMCP #AIEngineering #AIAgents #SoftwareArchitecture #OpenAI Companion devpost.com