<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Modular Technology Group</title>
	<atom:link href="https://modtechgroup.com/feed/" rel="self" type="application/rss+xml" />
	<link>https://modtechgroup.com/</link>
	<description></description>
	<lastBuildDate>Wed, 29 Jul 2026 12:16:09 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.0.2</generator>
	<item>
		<title>What &#8216;private AI&#8217; actually means</title>
		<link>https://modtechgroup.com/what-private-ai-actually-means/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=what-private-ai-actually-means</link>
		
		<dc:creator><![CDATA[Arthur]]></dc:creator>
		<pubDate>Wed, 29 Jul 2026 12:16:09 +0000</pubDate>
				<category><![CDATA[Modular]]></category>
		<guid isPermaLink="false">https://modtechgroup.com/what-private-ai-actually-means/</guid>

					<description><![CDATA[<p>What "private AI" actually meansModular house · Explainer, Private AI · ~4 min readPrivate AI is your models and your agents running on infrastructure you control, with your data staying inside boundaries you set. The name is about who holds the keys. It says nothing about how capable the tools are. Your data, your rules,  [Read more...]</p>
<p>The post <a href="https://modtechgroup.com/what-private-ai-actually-means/">What &#8216;private AI&#8217; actually means</a> appeared first on <a href="https://modtechgroup.com">Modular Technology Group</a>.</p>
]]></description>
										<content:encoded><![CDATA[<h1>What &#8220;private AI&#8221; actually means</h1>
<p>Modular house · Explainer, Private AI · ~4 min read</p>
<p>Private AI is your models and your agents running on infrastructure you control, with your data staying inside boundaries you set. The name is about who holds the keys. It says nothing about how capable the tools are. Your data, your rules, from dirt to desktop, one partner from the hardware to the interface. Here is what that looks like in practice, and what it does not require you to give up.</p>
<h2>The common misconception</h2>
<p>Say &#8220;private AI&#8221; in a planning meeting and watch what people picture. Usually something small. A stripped-down model wheezing away on a spare laptop, good for a demo and not much else, the discount version of the tools everyone already uses at work. That picture is years out of date. A private workspace runs current open-weight models, several of them, on real hardware in a real facility, and you pick the one that fits the job. Privacy and capability are two separate questions. Answering one does not cost you the other.</p>
<p>The confusion is easy to forgive, because the line got blurry on purpose. The public tools sell a business tier with a setting that promises your data will not be used for training, and that promise is worth something. Read where it lives, though. It lives in a contract, and a contract is a thing both parties can revisit. Meanwhile your prompts still leave your building, cross the open internet, and land on servers you have never seen, in a jurisdiction you did not pick. The setting governs what the vendor says it will do with your data. Where your data goes stays the same.</p>
<p>Private AI closes that gap by removing the trust exercise entirely. Data cannot leak out of a boundary when the computing happens inside the boundary. There is nothing to opt out of. A setting is a promise. A boundary is a fact.</p>
<h2>The three things that make AI private</h2>
<p>Start with where the model runs, because everything else follows from it. When you ask a public tool a question, the question travels. It leaves your network, crosses the internet, and gets processed in a data center you cannot name, alongside everyone else&#8217;s traffic. Private AI turns that around: instead of sending your data to the model, you bring the model to your data. For our clients that means an isolated environment in a US-based, FedRAMP-certified data center, and for organizations that want it, dedicated hardware on their own premises. The question gets answered a few racks away from the data it draws on. Sometimes a few feet.</p>
<p>Then the data itself, which covers a lot more than chat history. The real value of a workspace shows up when you connect it to your documents. Contracts, research, case files, the institutional memory of the whole company. That happens through retrieval: the model reads the relevant pieces of your knowledge base at query time and grounds its answer in them. In a private environment, the documents, the index built from them, and the retrieval itself all stay inside the same boundary. Nothing is used to train anyone&#8217;s model. Nothing is retained by a third party, because there is no third party in the room.</p>
<p>And the part that gets discussed least while mattering most: who draws the boundary. Who can see which documents. Whether the system can reach the internet at all. What gets encrypted, what gets backed up, and what happens to all of it if you decide to leave. On public platforms those rules are set by the vendor and adjusted at the vendor&#8217;s discretion. In a private environment they are yours to set: access controls that separate the legal team&#8217;s vault from marketing&#8217;s, an air gap where the work demands one, and an exit that is just your data, in usable formats, walking out the door with you. That last one tells you who really owned the environment. If leaving is easy, you did.</p>
<h2>What you keep</h2>
<p>Everything your team actually liked about the public tools. The chat interface, the drafting and summarizing, the document questions, the code help. Those run just as well on infrastructure you control, through a clean web interface with access to multiple models, so people pick the model that fits the task instead of the one a vendor is promoting this quarter. Good UX is a software problem, and the software exists. Nobody has to learn to love a command line.</p>
<p>You also keep room to grow, because private is a spectrum, and you do not have to buy the far end of it on day one. Our Wildcat tier is the entry point: an isolated environment on shared infrastructure, in the same FedRAMP facility, under the same US jurisdiction as everything above it. Panther steps up to dedicated infrastructure, with file vaults, role-based access controls, and enhanced encryption for teams handling sensitive documents. Grizzly is the top: fully dedicated hardware, zero-trust architecture, an optional air gap, and the option to run it on your own premises instead of ours. The boundary tightens as you climb. What never changes is the jurisdiction. Your data stays in the United States whether it sits in our facility or in your building.</p>
<p>And you keep one accountable partner across the whole stack. The facility, with its redundant power and its guarded doors. The hardware in the racks. The models, kept patched and current, swapped for better ones as better ones ship. The interface your people log into every morning. When something needs attention, there is no seam between a cloud vendor, a hosting vendor, and an AI vendor for the problem to fall through. You hold the deed, and there is one number to call.</p>
<p>Your data, your rules.</p>
<p>The post <a href="https://modtechgroup.com/what-private-ai-actually-means/">What &#8216;private AI&#8217; actually means</a> appeared first on <a href="https://modtechgroup.com">Modular Technology Group</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>The hidden cost of renting your AI</title>
		<link>https://modtechgroup.com/hidden-cost-of-renting-your-ai/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=hidden-cost-of-renting-your-ai</link>
		
		<dc:creator><![CDATA[Arthur]]></dc:creator>
		<pubDate>Mon, 27 Jul 2026 23:16:05 +0000</pubDate>
				<category><![CDATA[Modular]]></category>
		<guid isPermaLink="false">https://modtechgroup.com/hidden-cost-of-renting-your-ai/</guid>

					<description><![CDATA[<p>The hidden cost of renting your AIFrom the Desk of Cale · Cost, Fixed-cost AI · ~5 min readPer-token pricing is a great way to start and a hard way to budget. The variable cost of someone else's AI infrastructure isn't actually variable to you. You just absorb the outputs: pricing changes, capacity limits, decisions  [Read more...]</p>
<p>The post <a href="https://modtechgroup.com/hidden-cost-of-renting-your-ai/">The hidden cost of renting your AI</a> appeared first on <a href="https://modtechgroup.com">Modular Technology Group</a>.</p>
]]></description>
										<content:encoded><![CDATA[<h1>The hidden cost of renting your AI</h1>
<p>From the Desk of Cale · Cost, Fixed-cost AI · ~5 min read</p>
<p>Per-token pricing is a great way to start and a hard way to budget. The variable cost of someone else&#8217;s AI infrastructure isn&#8217;t actually variable to you. You just absorb the outputs: pricing changes, capacity limits, decisions made at a scale you&#8217;ll never see. Fixed, right-sized private infrastructure flips that. You know what you&#8217;re paying, and you know why.</p>
<h2>The day-one price is not the real price</h2>
<p>I&#8217;ve been thinking about taxi meters. A taxi is the right call for a trip to the airport, and nobody argues with the meter for one ride. But if you found yourself taking that same taxi to work every morning, you&#8217;d start doing math on a car payment before the end of the month. Per-token AI is a taxi meter, and a lot of companies are now commuting in it. Nobody chose the commute, either. The meter was already running when the habit formed.</p>
<p>The day-one price looks great because day-one usage is tiny. A handful of questions, a summarized document or two. Fractions of a cent each. Then the tool turns out to be useful, which is the whole point, and useful tools get used. People stop asking one question and start pasting in whole contracts. Someone wires it into a workflow that runs on every support ticket. Then come the agents, and an agent doesn&#8217;t make one call per task. It reads the document, checks its own work, calls a tool, reads the result, tries again. Every one of those steps is a metered ride. The per-token price never moved. Your consumption did, while everyone was busy being productive.</p>
<p>I sat with a client this spring and put a year of their AI invoices side by side on one screen. That was the entire exercise. No spreadsheet wizardry, no consultants. The line went one direction, and nobody in the room could name a month where anyone decided to spend more. That&#8217;s the tell. Metered costs don&#8217;t get decided. They accumulate.</p>
<h2>Variable to them, fixed to you (and vice versa)</h2>
<p>Here&#8217;s the part that took me embarrassingly long to see clearly. The provider&#8217;s costs are mostly fixed. The data centers are already built and the payroll is already set. What&#8217;s variable, to them, is you. Metered pricing is how they convert their fixed cost into your variable one. That&#8217;s a rational move on their side of the table. It&#8217;s just worth noticing which side of the table you&#8217;re on. When a business absorbs volatility, it charges for the service. When it passes volatility through, you&#8217;re the one providing that service, and nobody&#8217;s paying you for it.</p>
<p>So when their world shifts, the shift gets passed through. A new model generation lands and the price per token changes. Demand spikes and rate limits show up at exactly your busy hour. An older model gets retired, and the workflow your team spent a quarter tuning now runs on something that behaves differently. There&#8217;s no villain in any of that. A business planning for millions of customers makes ordinary capacity decisions, and you&#8217;re one line in the plan. You don&#8217;t get a vote. You get an email.</p>
<p>Now put yourself in the budget meeting. Finance asks what AI will cost next year. The honest answer under metered pricing is &#8220;it depends on how much people use it,&#8221; which is another way of saying the better it works, the less we can predict. That&#8217;s a strange incentive to hand a mid-sized team. Success becomes a cost overrun. I&#8217;ve watched a manager quietly discourage adoption of a tool his own company was paying for, because every new enthusiastic user made his forecast worse. Every budget is a guess, but this one is a guess about other people&#8217;s guesses.</p>
<h2>What predictable looks like</h2>
<p>Predictable starts with a boring question: what do you actually run? Not someday, today. Count the real workloads. The document review, the drafting, the internal search, the two or three automations that matter. Most teams find the list is shorter than they feared and steadier than the invoices implied. That steadiness is the asset. You size for your own team and the work it actually does. The whole internet is somebody else&#8217;s capacity problem. Once you know the workload, you can size the hardware to it, and once the hardware is sized, the cost is flat. You know the number in January and it&#8217;s still the number in October.</p>
<p>That&#8217;s the shape of what we build at Modular Technology Group. Our Private AI Workspaces start with Wildcat on shared infrastructure, step up to Panther on dedicated infrastructure, and top out at Grizzly on fully dedicated hardware, which you can host with us or stand up in your own building. The hosted tiers all live in a US-based, FedRAMP-certified data center. Every one of them bills as a flat monthly number. No per-token billing anywhere. We run our own work on the same stack, so when the meter argument comes up, we&#8217;re not speculating. We live on the fixed side of it. Right-sizing is a conversation. Some teams land on shared infrastructure and stay there happily. Some need the hardware where they can see it. Either way the number gets chosen, and you&#8217;re in the room when it happens.</p>
<p>The total-cost math is worth saying plainly. Owning can look more expensive on day one, the way a car payment looks worse than one cab fare. But a fixed cost changes the direction of every incentive after that. Under a meter, each new user is a liability. On infrastructure you own, each new user makes every task cheaper, because the same monthly number is now doing more work. You quit rationing the tool and start pushing it. Adoption stops being a cost problem and turns back into what it should have been, which is a productivity story. And because the models sit behind an interface you control, no single provider&#8217;s pricing decision can reach into your budget. When a better open model ships, it slots in. The bill doesn&#8217;t notice.</p>
<p>Here&#8217;s something you can do this week, and it costs nothing. Pull your last twelve months of AI invoices and put them in one place. Ask two questions. What happens to this line if usage doubles, and who decided the current number? If the answers are &#8220;it doubles&#8221; and &#8220;nobody,&#8221; you&#8217;re renting, and now you know what the rent really is. Then you get to decide what the number should be instead. Your data, your rules, from dirt to desktop.</p>
<p>If your AI line item has quit behaving, or you just want a second set of eyes on the own-versus-meter math, I&#8217;m always glad to compare notes. No pitch. Bring the invoices.</p>
<p>Own it, don&#8217;t rent it. Your data, your rules.</p>
<p>The post <a href="https://modtechgroup.com/hidden-cost-of-renting-your-ai/">The hidden cost of renting your AI</a> appeared first on <a href="https://modtechgroup.com">Modular Technology Group</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>The Modular Briefing, Episode 8: Who Gave It the Keys?</title>
		<link>https://modtechgroup.com/who-gave-it-the-keys/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=who-gave-it-the-keys</link>
		
		<dc:creator><![CDATA[Arthur]]></dc:creator>
		<pubDate>Sun, 26 Jul 2026 13:31:30 +0000</pubDate>
				<category><![CDATA[Agentic AI]]></category>
		<category><![CDATA[Podcast]]></category>
		<category><![CDATA[agentic]]></category>
		<category><![CDATA[agents]]></category>
		<category><![CDATA[AI guardrails]]></category>
		<category><![CDATA[oversight]]></category>
		<guid isPermaLink="false">https://modtechgroup.com/who-gave-it-the-keys/</guid>

					<description><![CDATA[<p>One link can create an agent with your connectors attached. Who hands out that authority, who watches it, and who decides where it runs.</p>
<p>The post <a href="https://modtechgroup.com/who-gave-it-the-keys/">The Modular Briefing, Episode 8: Who Gave It the Keys?</a> appeared first on <a href="https://modtechgroup.com">Modular Technology Group</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<div style="margin:0 0 24px;">
<audio controls preload="metadata" style="width:100%;" src="https://assets.modtechgroup.com/podcast/audio/modular-briefing-ep08.mp3"></audio>
<p style="font-size:16px;color:#ffffff;margin:6px 0 0;">Episode 8 &middot; 4:40 &middot; <a href="https://assets.modtechgroup.com/podcast/audio/modular-briefing-ep08.mp3">Download MP3</a> &middot; <a href="https://assets.modtechgroup.com/podcast/feed.xml">RSS</a> &middot; <a href="#transcript">Transcript</a></p>
</div>



<p class="wp-block-paragraph">An AI agent is not a feature you switch on. It is an actor with authority. This week: a single link that could build a working agent inside somebody&#8217;s workspace with every connector attached and approvals switched off, an attacker who ran an agent unattended inside a national finance ministry, and a government that stopped a billion-euro cloud procurement to ask where its data lives. Three sizes of the same question.</p>



<h2 class="has-text-color wp-block-heading" style="color:#ffffff">In this episode</h2>



<ul class="wp-block-list">
<li><strong>A link that built an agent.</strong> Researchers showed one crafted URL could create a working agent inside a logged-in account, attach every connected mailbox and file store, set approvals to never ask, and put it on an hourly schedule. Patched on June 8, with no reported exploitation in the wild. The open question is not the bug. It is how casually agent authority gets handed out.</li>
<li><strong>An agent nobody was supervising.</strong> A security firm reports an attacker running an off-the-shelf agent with its approval mode disabled for post-exploitation work inside a national finance ministry. One firm is the only public source and the ministry has not confirmed it. The transferable lesson is that the useful setting was the one that removed the human.</li>
<li><strong>A government asking where its data lives.</strong> Ireland&#8217;s Office of Government Procurement cancelled a cloud framework competition over digital sovereignty. Jurisdiction moved from a slide in a security review to a reason to stop a contract.</li>
</ul>



<h2 class="has-text-color wp-block-heading" style="color:#ffffff">Sources</h2>



<ul class="wp-block-list">
<li><a href="https://thehackernews.com/2026/07/chatgpt-agentforger-flaw-could-deploy.html" rel="nofollow">The Hacker News, July 24, 2026 &mdash; agent-builder flaw</a></li>
<li><a href="https://thehackernews.com/2026/07/hacker-runs-hermes-ai-agent-unattended.html" rel="nofollow">The Hacker News, July 24, 2026 &mdash; unattended agent at a national finance ministry</a></li>
<li><a href="https://www.theregister.com/public-sector/2026/07/22/ireland-stalls-1b-microsoft-tender-amid-digital-sovereignty-questions/5276149" rel="nofollow">The Register, July 22, 2026 &mdash; Ireland stalls cloud tender over digital sovereignty</a></li>
</ul>



<h2 class="has-text-color wp-block-heading" style="color:#ffffff">AI voice disclosure</h2>



<p class="wp-block-paragraph">Laura and Arthur are AI-generated voices, produced locally on Modular&#8217;s own hardware. Story selection, reporting, and fact-checking are done by the Modular Technology Group team. Your data, your rules applies to our own production too.</p>



<details id="transcript" style="margin:20px 0;border:1px solid #E1E8ED;border-radius:10px;padding:14px 18px;">
<summary style="cursor:pointer;font-weight:700;">Full transcript</summary>
<div style="margin-top:12px;font-size:15px;line-height:1.6;">
<p><strong>Laura:</strong> Welcome to The Modular Briefing, the show that cuts through the AI noise and tells you what it actually means for your business. I&#8217;m Laura.</p>
<p><strong>Arthur:</strong> And I&#8217;m Arthur. Three stories today, and one idea underneath all of them. An AI agent is an actor with authority. So today, who hands out that authority. Who is watching it. And who gets to decide where any of it runs.</p>
<p><strong>Laura:</strong> Story one. Security researchers showed that on a widely used AI platform, one crafted link was enough to build a working agent inside somebody&#8217;s account. The victim only had to be logged in and click. The agent came up from a template, attached every mailbox and file store that account had already connected, set its approval prompts to never ask, and put itself on an hourly schedule. To be fair to the vendor, they fixed it on the eighth of June, and nobody has reported it being used in the wild.</p>
<p><strong>Arthur:</strong> Picture that at the desk. Somebody in accounting clicks a link in an email, and forty minutes later there is a thing in their workspace reading the mailbox on a timer, with permission to act, and no human ever approved it. The patch closes that one door. It does not answer the question the story asks, which is how casually agent authority gets handed out in the first place. In most companies right now, anybody with a login can create an agent, wire it to real data, and nobody can name who owns it. That is the part you can fix this week. Give every agent a boundary and a person accountable for it, in a workspace your company actually owns rather than one you rent by the seat. Your data, your rules.</p>
<p><strong>Laura:</strong> Story two is what that looks like when nobody is watching. A security firm reports that an attacker ran an off-the-shelf AI agent inside a national finance ministry, using it for the messy work after a break-in, with the agent&#8217;s approval mode switched off so it would not stop to ask. Researchers found the operator&#8217;s own toolkit sitting in an exposed directory, around five hundred and eighty-five files. Worth saying plainly: one firm is the only public source, the ministry has not confirmed any of it, and their national cyber team was notified in mid July.</p>
<p><strong>Arthur:</strong> Take the caveat seriously and the lesson still stands, because the interesting detail is not the break-in. It is that the useful setting was the one that turns the human out of the loop. That setting exists on the tools your own team uses. If somebody on your staff can disable an approval gate to move faster, then the gate was decoration. Real oversight means the boundary sits outside the agent, in infrastructure you control, and it means you have an off switch that works even when the agent is mid-task. Your data, your rules, and that includes the AI working on it. Your AI, your rules.</p>
<p><strong>Laura:</strong> Story three moves the question up a level. Ireland&#8217;s government procurement office just cancelled a major cloud framework competition after concerns were raised about digital sovereignty. For scale, the framework it would have replaced is capped at three hundred and fifty million euro and runs to September of twenty twenty-seven, and an opposition deputy put the replacement somewhere between seven hundred and fifty million and one billion euro.</p>
<p><strong>Arthur:</strong> Here is why that matters to a twelve-person firm in Kentucky. A government just treated the question of where its data lives, and whose law reaches it, as a reason to stop a procurement in its tracks. That question used to be a slide in a security review. Now it moves contracts. If a national government is willing to pause and ask it, it is a fair question for you to ask about the systems running your client files. And you get better answers when the stack is yours: infrastructure you own, in a US-based FedRAMP facility, at a fixed monthly cost instead of a meter you cannot see, one stack from dirt to desktop. Your data, your rules.</p>
<p><strong>Laura:</strong> That is the thread today. A link that built an agent. An agent nobody was supervising. A government that stopped a contract to ask where its data lives. Same question at three different sizes.</p>
<p><strong>Arthur:</strong> If your team is starting to point AI agents at real systems and you are not sure who owns them or what they are allowed to touch, we would genuinely like to compare notes. No hard pitch. Head to modtechgroup dot com slash consultation and book a conversation with the Modular team. We will help you figure out where your data lives, what it really costs, and what your options are.</p>
<p><strong>Laura:</strong> Thanks for spending a few minutes with us.</p>
<p><strong>Laura:</strong> This has been The Modular Briefing.</p>
<p><strong>Arthur:</strong> Your data, your rules.</p>
<p><strong>Laura:</strong> We will see you next time.</p>
</div>
</details>
<script>
(function(){
  function openTranscript(){
    var d=document.getElementById('transcript');
    if(d&&location.hash==='#transcript'){d.open=true;d.scrollIntoView();}
  }
  window.addEventListener('hashchange',openTranscript);
  document.addEventListener('DOMContentLoaded',openTranscript);
  document.addEventListener('click',function(e){
    var a=e.target.closest&&e.target.closest('a[href="#transcript"]');
    if(a){var d=document.getElementById('transcript');if(d){d.open=true;}}
  });
})();
</script>



<p class="wp-block-paragraph">If your team is pointing AI agents at real systems and nobody can say who owns them or what they may touch, <a href="https://modtechgroup.com/consultation/?utm_source=podcast&amp;utm_medium=episode&amp;utm_campaign=ep08">book a conversation</a>.</p>



<p class="wp-block-paragraph"><strong>Your data, your rules.</strong></p>
<p>The post <a href="https://modtechgroup.com/who-gave-it-the-keys/">The Modular Briefing, Episode 8: Who Gave It the Keys?</a> appeared first on <a href="https://modtechgroup.com">Modular Technology Group</a>.</p>
]]></content:encoded>
					
		
		<enclosure url="https://assets.modtechgroup.com/podcast/audio/modular-briefing-ep08.mp3" length="4482238" type="audio/mpeg" />

			</item>
		<item>
		<title>The compliance gate paused. Your duty to your data did not.</title>
		<link>https://modtechgroup.com/compliance-gate-paused-duty-to-data/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=compliance-gate-paused-duty-to-data</link>
		
		<dc:creator><![CDATA[Arthur]]></dc:creator>
		<pubDate>Fri, 24 Jul 2026 13:05:30 +0000</pubDate>
				<category><![CDATA[Compliance]]></category>
		<guid isPermaLink="false">https://modtechgroup.com/compliance-gate-paused-duty-to-data/</guid>

					<description><![CDATA[<p>Compliance · CMMC, Private AI, Governance · ~6 min readOn July 13, 2026, the Department of War announced the immediate suspension of CMMC Phase II, the third-party certification requirement that had been scheduled to start appearing in defense contracts on November 10, 2026, roughly four months out (DefenseScoop). Read the announcement closely: the suspension covers  [Read more...]</p>
<p>The post <a href="https://modtechgroup.com/compliance-gate-paused-duty-to-data/">The compliance gate paused. Your duty to your data did not.</a> appeared first on <a href="https://modtechgroup.com">Modular Technology Group</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p><em>Compliance · CMMC, Private AI, Governance · ~6 min read</em></p>
<p>On July 13, 2026, the Department of War announced the immediate suspension of CMMC Phase II, the third-party certification requirement that had been scheduled to start appearing in defense contracts on November 10, 2026, roughly four months out (<a href="https://defensescoop.com/2026/07/13/dod-halts-cmmc-cybersecurity-requirements-phase-2/">DefenseScoop</a>). Read the announcement closely: the suspension covers only the certification gate. DFARS 7012 still applies. The NIST 800-171 self-assessment still applies. For anyone bringing AI into the defense supply chain, that means less paperwork sitting on top of the same responsibility. Piping that data into someone else&#8217;s cloud model is still a risk you own.</p>
<p>What the suspension actually covers</p>
<p>Phase II is the part of CMMC where an outside assessor certifies you, meaning Level 2 assessments conducted by a C3PAO and the Level 3 assessments above them. That requirement is now suspended while a new CMMC Reform Task Force runs a 60-day review, and Phases 3 and 4 are frozen along with the program&#8217;s future implementation milestones. The stated reason is cost. The department wants to cut compliance expense and bureaucratic burden as part of a wider effort to streamline defense acquisition, and for a small contractor that is genuinely good news.</p>
<p>The review leaves the standing rules in place. DFARS clause 252.204-7012, which requires you to safeguard covered defense information and report cyber incidents, remains in force. So does the CMMC Phase I self-assessment: you still complete a NIST SP 800-171 self-assessment and upload your score to the DoD&#8217;s SPRS system. An inaccurate score can still create liability under the False Claims Act. The only thing removed from the calendar is the assessor&#8217;s visit. Every rule that assessor would have checked remains on the books.</p>
<p>What the pause means for how you handle CUI</p>
<p>The obligation follows the data, and it always has. Controlled unclassified information, CUI, carried the same obligations on July 14 that it carried on July 12, whether or not anyone is scheduled to check your work. Now consider where CUI actually goes when someone on your team pastes a spec or a contract excerpt into a public AI chat tool. It leaves your network and lands on infrastructure you cannot inspect, under retention terms you never negotiated. DFARS 7012 still expects you to report incidents involving that information, and you cannot report what you cannot see.</p>
<p>There is also the score you already attested. Your SPRS submission describes an environment where CUI stays inside defined controls. If CUI is quietly flowing into public tools, that submission stops being true, and the False Claims Act does not pause for a task force. The right response is simply to know where your data goes, and that is a thing you get to decide.</p>
<p>The private-AI answer for defense work</p>
<p>The same move that satisfies the standing rules also unlocks the AI you wanted in the first place: run the model inside a boundary you control. That can be your own environment, or a private one a partner builds and operates for you. When the model lives where the data lives, CUI never crosses into a system you cannot account for. This maps directly onto what survived the review. 7012 asks you to safeguard covered defense information; a boundary you control is the safeguard. The 800-171 self-assessment asks you to score your environment honestly; a contained environment is one you can score honestly and keep scoring honestly, review or no review.</p>
<p>You can move before the task force reports, and you can bring in help to do it. Modular Technology Group builds and runs private AI for contractors in exactly this position, and we own the whole stack. Your data, your rules, from dirt to desktop. Use the pause to get your AI inside your boundary, so that whatever comes back from the review, you are already standing where the rules point.</p>
<p>Your data, your rules.</p>
<p>The post <a href="https://modtechgroup.com/compliance-gate-paused-duty-to-data/">The compliance gate paused. Your duty to your data did not.</a> appeared first on <a href="https://modtechgroup.com">Modular Technology Group</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>AI needs a home, not a hotel</title>
		<link>https://modtechgroup.com/ai-needs-a-home-not-a-hotel/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=ai-needs-a-home-not-a-hotel</link>
		
		<dc:creator><![CDATA[Arthur]]></dc:creator>
		<pubDate>Thu, 23 Jul 2026 14:46:15 +0000</pubDate>
				<category><![CDATA[Modular]]></category>
		<guid isPermaLink="false">https://modtechgroup.com/ai-needs-a-home-not-a-hotel/</guid>

					<description><![CDATA[<p>From the Desk of Cale · Private AI · ~4 min readMost companies did not decide to rent their AI. It happened one login at a time. A hotel is fine for a night, but you do not own the room, you cannot change the locks, and the nightly rate is whatever they say it  [Read more...]</p>
<p>The post <a href="https://modtechgroup.com/ai-needs-a-home-not-a-hotel/">AI needs a home, not a hotel</a> appeared first on <a href="https://modtechgroup.com">Modular Technology Group</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p><em>From the Desk of Cale · Private AI · ~4 min read</em></p>
<p>Most companies did not decide to rent their AI. It happened one login at a time. A hotel is fine for a night, but you do not own the room, you cannot change the locks, and the nightly rate is whatever they say it is. Your business AI deserves a home you own, where your data stays yours and the bill does not surprise you.</p>
<p>How renting happens by accident</p>
<p>Nobody signs a lease with the whole company in mind. Someone in marketing starts a free trial to draft campaign copy. An engineer signs up with a work email because the tool saves an hour a day. Neither asks permission, because why would they? It is a free trial. Six months later there are a dozen seats spread across four departments, nobody owns any of them, and nobody can say what has been pasted into which chat window. There is no boundary because nobody ever drew one.</p>
<p>The moment you notice is usually mundane. A renewal notice arrives at a new rate, or a client&#8217;s security questionnaire asks where their information lives and the honest answer takes a week to assemble. So you decide to consolidate, and you discover that the prompts, the saved conversations, the custom workflows, the habits your team built are all inside someone else&#8217;s building. You can walk out, but you leave the furniture. That is the tell that you have been renting all along: leaving costs more than staying, and the landlord knows it.</p>
<p>What &#8220;owning the room&#8221; actually buys you</p>
<p>Start with the question your clients are already asking you: where does our data live? When you own the room, the answer is one sentence. It lives here, inside a boundary we control, and it does not feed anyone else&#8217;s model. That answer shortens security reviews and calms auditors, and it has the advantage of being simply true. No contract exhibit required.</p>
<p>The economics change too. Renting means metered billing, per seat and per token, at a rate set by someone who knows you cannot easily leave. Owning means the cost is mostly fixed. You know what the hardware costs and what running it costs, and the bill in month eighteen looks like the bill in month three. Budgeting becomes boring, which is exactly what budgeting should be. There is a quieter benefit underneath both of those: because the models sit behind an interface you control, you can change them without changing anything else. When a better open model ships, and they ship constantly now, you swap it in over a weekend. No migration project, no new vendor negotiation. The room stays the same. Only the furniture improves.</p>
<p>What it does not require you to give up</p>
<p>The usual objection is that owning means going without. It does not. The tools your team already likes (the chat interface, the document work, the code assistance, the meeting summaries) run just as well on infrastructure you control. Good UX is a software problem, and the software exists. Your people should not notice a difference, except that the question of where the data goes finally has an answer.</p>
<p>Owning also does not mean building from scratch. You would not pour your own foundation to own a house; you would hire a builder and hold the deed. The same trade exists here. A partner can stand up the hardware, run the models, keep everything patched and current, and hand you the keys. That is the work we do at Modular: your data, your rules, from dirt to desktop. You hold the deed. We keep the lights on.</p>
<p>Your data, your rules.</p>
<p>The post <a href="https://modtechgroup.com/ai-needs-a-home-not-a-hotel/">AI needs a home, not a hotel</a> appeared first on <a href="https://modtechgroup.com">Modular Technology Group</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>When Your Team Ships Work They Can&#8217;t Explain</title>
		<link>https://modtechgroup.com/when-your-team-ships-work-they-cant-explain/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=when-your-team-ships-work-they-cant-explain</link>
		
		<dc:creator><![CDATA[Arthur]]></dc:creator>
		<pubDate>Wed, 15 Jul 2026 14:02:18 +0000</pubDate>
				<category><![CDATA[AI Governance]]></category>
		<category><![CDATA[fcaio]]></category>
		<category><![CDATA[oversight]]></category>
		<category><![CDATA[shadow ai]]></category>
		<guid isPermaLink="false">https://modtechgroup.com/when-your-team-ships-work-they-cant-explain/</guid>

					<description><![CDATA[<p>Heavy AI users are shipping work they don't fully understand. The fix isn't banning AI, it's governing it: an accountable owner, sanctioned tools, and training.</p>
<p>The post <a href="https://modtechgroup.com/when-your-team-ships-work-they-cant-explain/">When Your Team Ships Work They Can&#8217;t Explain</a> appeared first on <a href="https://modtechgroup.com">Modular Technology Group</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p>A new report making the rounds this week found something most managers already suspected but had not measured: the heaviest AI users are the ones most likely to hand in work they cannot fully explain. They are fast. The output looks polished. And when you ask how a number was derived or why a clause reads the way it does, the answer is some version of &#8220;the AI wrote it.&#8221;</p>
<p>That is not a story about lazy employees. It is a story about missing guardrails.</p>
<h2>The real risk is ownership, not the tool</h2>
<p>When someone submits work they do not understand, the organization has quietly taken on a liability it cannot see. A wrong figure in a board deck, an unsupported claim in a client memo, a compliance answer that sounds right and is not: each of these is now traveling under your name, and no one in the building can defend it. The tool did not create that exposure. The absence of an owner did.</p>
<p>This is the pattern behind most shadow AI. People adopt AI because it helps, not because they were reckless. They paste a contract into a public chatbot to summarize it. They let a model draft the analysis and move on. Every one of those choices is reasonable in isolation. Added up across a department, with no policy and no review, they become an accountability gap that widens every week.</p>
<h2>Banning the tool is the wrong instinct</h2>
<p>The reflex in regulated and high-stakes work is to lock AI down. Block the sites, forbid the tools, wait for the risk to pass. It never does. The work still needs doing, the tools are still faster, and the usage simply moves somewhere you cannot see. A ban does not remove the risk. It removes your visibility into it.</p>
<p>The better move is to govern the tool the same way you govern any other capability that touches sensitive work. Name an owner. Set the boundaries. Give people sanctioned tools that are good enough that they stop reaching for the unsanctioned ones. Your AI, your rules, the same principle you already apply to your data.</p>
<h2>What governance looks like in practice</h2>
<p>Governance sounds heavy. In a mid-sized company it is closer to a short list of decisions made once and enforced consistently:</p>
<ul>
<li><strong>An accountable owner.</strong> Someone whose job is to say what AI is allowed to touch, and to answer for it. This is the core of the fractional Chief AI Officer role: senior judgment on AI risk without a full-time executive hire.</li>
<li><strong>Sanctioned tools with real boundaries.</strong> A private, governed workspace where your data stays inside your walls and never trains someone else&#8217;s model. When the approved option is genuinely useful, shadow AI loses its pull.</li>
<li><strong>Oversight that fits the work.</strong> Higher-stakes output gets a human check before it ships. Routine work moves fast. Match the level of review to the level of consequence, not the other way around.</li>
<li><strong>Training that closes the gap.</strong> People should be able to explain what the AI did for them. That is a skill, and it is teachable.</li>
</ul>
<p>None of this requires banning anything. It requires deciding who is responsible before the work goes out the door, not after it comes back wrong.</p>
<h2>The through-line</h2>
<p>Every organization is going to run on more AI next year than it does today. The question is not whether your team uses it. They already do. The question is whether the work it produces is something you can stand behind.</p>
<p>That comes down to ownership. Own the tools, own the data, own the decision about who is accountable. A duty of care, not a ban.</p>
<p>If you are not sure who owns AI risk in your organization, that is the first thing worth fixing. It is exactly the problem a fractional Chief AI Officer is built to solve, and it is where we usually start.</p>
<p><a href="https://modtechgroup.com/consultation/"><strong>Talk to us about AI governance &rarr;</strong></a></p>
<p><em>Your data, your rules. Modular Technology Group helps regulated and mid-sized teams run private AI on infrastructure they own, from dirt to desktop.</em></p>
<p>The post <a href="https://modtechgroup.com/when-your-team-ships-work-they-cant-explain/">When Your Team Ships Work They Can&#8217;t Explain</a> appeared first on <a href="https://modtechgroup.com">Modular Technology Group</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>When the Token Bill Comes Due: What Uber and Microsoft Just Taught the Rest of Us About Renting Intelligence</title>
		<link>https://modtechgroup.com/when-the-token-bill-comes-due-what-uber-and-microsoft-just-taught-the-rest-of-us-about-renting-intelligence/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=when-the-token-bill-comes-due-what-uber-and-microsoft-just-taught-the-rest-of-us-about-renting-intelligence</link>
		
		<dc:creator><![CDATA[Cale Hollingsworth]]></dc:creator>
		<pubDate>Tue, 02 Jun 2026 16:51:30 +0000</pubDate>
				<category><![CDATA[AI Infrastructure]]></category>
		<category><![CDATA[cost]]></category>
		<category><![CDATA[fixed cost ai]]></category>
		<category><![CDATA[own vs rent]]></category>
		<guid isPermaLink="false">https://modtechgroup.com/?p=5773</guid>

					<description><![CDATA[<p>In late May, Goldman Sachs put a number on something a lot of operators have been feeling in their gut for months. Agentic AI, the bank projects, could push token demand up by more than 24 times in the next few years. Read that again. Not 24 percent. Twenty-four times. If your AI runs on  [Read more...]</p>
<p>The post <a href="https://modtechgroup.com/when-the-token-bill-comes-due-what-uber-and-microsoft-just-taught-the-rest-of-us-about-renting-intelligence/">When the Token Bill Comes Due: What Uber and Microsoft Just Taught the Rest of Us About Renting Intelligence</a> appeared first on <a href="https://modtechgroup.com">Modular Technology Group</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">In late May, Goldman Sachs put a number on something a lot of operators have been feeling in their gut for months. Agentic AI, the bank projects, could push token demand up by more than 24 times in the next few years. Read that again. Not 24 percent. Twenty-four times.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">If your AI runs on someone else&#8217;s meter, that forecast is not a growth story. It&#8217;s an invoice you haven&#8217;t opened yet.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">And the companies you&#8217;d expect to absorb that hit better than anyone are the ones flinching first. Uber reportedly burned through its entire 2026 AI budget in a matter of months. Uber&#8217;s CTO went public with it, and the operations chief, Andrew Macdonald, told Business Insider something even more damning than the overspend: after talking to his senior engineers, he couldn&#8217;t find a clear line between how many tokens the company was burning and how many features customers actually got. More than 80 percent of Uber&#8217;s engineers were using agentic tools. Over 60 percent of the code was AI-generated. And it still wasn&#8217;t worth what they were paying for it.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">Microsoft, meanwhile, started pulling its own developers off a third-party coding assistant and moving them onto an in-house tool, with a deadline that landed conveniently at the close of its fiscal year. The official line was consolidation. The timing told a different story. Microsoft also flipped one of its developer products to token-based billing because the cost of running it had ballooned.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">When the two companies that helped write the playbook for aggressive AI adoption are both quietly restructuring how they buy it, that&#8217;s not a blip. That&#8217;s the meter catching up with the marketing.</p>
<h2 class="text-text-100 mt-3 -mb-1 text-[1.125rem] font-bold">The math nobody put on the slide</h2>
<p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">For two years, the pitch for cloud AI has been simple: usage is cheap, it&#8217;ll only get cheaper, and you can scale infinitely. The first part was true for a chatbot answering one question at a time. It stops being true the moment you point an agent at a real workflow.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">A single agentic task can consume more than a thousand times the tokens of a one-shot chatbot query. Agents don&#8217;t ask once. They plan, call tools, check their own work, retry, and chain steps together, and every one of those steps is metered. Multiply that by a whole department running agents all day, then layer on the Goldman Sachs 24x demand curve, and the &#8220;it&#8217;ll get cheaper&#8221; story collapses under its own arithmetic.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">The numbers coming out of the industry have started to sound less like efficiency and more like a dare. Nvidia&#8217;s CEO said earlier this year that if one of his $500,000 engineers wasn&#8217;t burning at least $250,000 in tokens, he&#8217;d be worried. Airbnb&#8217;s CEO bragged that 60 percent of the company&#8217;s code is now AI-generated. One reported that 84 percent of its code was AI-written. A three-person team running an aggressive stack of agents managed to spend $1.3 million in tokens in a single month.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">Somewhere in there, the conversation quietly stopped being about results and started being about consumption for its own sake. Token usage became the brag, as if the size of the bill proved the value of the work. Uber just demonstrated, in public, that it doesn&#8217;t.</p>
<h2 class="text-text-100 mt-3 -mb-1 text-[1.125rem] font-bold">Consumption is not a strategy</h2>
<p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">Here&#8217;s the part that should land for any business owner watching this from the outside: the meter punishes exactly the usage you&#8217;re being told to chase.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">You are encouraged to put AI into everything, hand the agents more autonomy, let them run longer and reason harder. Every one of those instructions increases token consumption. So the more seriously you take the advice, the faster your costs compound, and they compound on a curve you don&#8217;t control and can&#8217;t predict. You find out what you owe after the work is done. That&#8217;s a brutal way to run a budget, and it&#8217;s an impossible way to run a small or mid-sized organization that needs to know its number before the quarter starts, not after.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">This is the question we ask clients to sit with before they sign anything: <em>what happens to our costs if our usage succeeds?</em> If the honest answer is &#8220;they go up in a way we can&#8217;t forecast,&#8221; then the platform isn&#8217;t priced for you to win. It&#8217;s priced for you to ration.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">And rationing is precisely what&#8217;s happening. The biggest names in tech are now teaching their people to use less of the very tools they spent a year telling everyone to use more of. If Microsoft and Uber can&#8217;t make consumption-based AI pencil out at their scale, the odds that a 40-person law firm or a boutique advisory shop will are not good.</p>
<h2 class="text-text-100 mt-3 -mb-1 text-[1.125rem] font-bold">The hardware cavalry isn&#8217;t coming in time</h2>
<p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">The usual reassurance is that better chips will rescue the economics. Next-generation inference hardware is genuinely more efficient, and the Goldman Sachs report leans on exactly that hope: cheaper tokens, usage keeps climbing, profits eventually follow.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">The timing doesn&#8217;t cooperate. The newest platforms are still rolling out, and the efficiency gains, real as they are, are years from deploying at the scale this demand curve requires. In the meantime, more than half of the data center projects planned around the current generation of hardware have reportedly been delayed or cancelled, choked by shortages of power and parts. The hyperscalers themselves have started stretching their hardware to run for six years instead of replacing it on the old cadence, which is hard to square with the promise of a dramatic efficiency leap every single year.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">So the demand is exploding now. The relief is theoretical and late. And the gap between the two gets paid for in your monthly bill.</p>
<h2 class="text-text-100 mt-3 -mb-1 text-[1.125rem] font-bold">There is another way to buy this</h2>
<p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">None of this is an argument against AI. We build our business on AI. It&#8217;s an argument against renting your intelligence by the drink from infrastructure you don&#8217;t control, priced on a model designed to climb.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">At Modular Technology Group, we made a different bet, and the news this month is the reason we made it. Modular runs private AI on infrastructure we own, in a US-based FedRAMP data center, at a fixed monthly price. No per-token billing. No per-query billing. No consumption meter quietly compounding in the background while your agents do exactly what you asked them to do.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">When the AI runs on hardware you control, the equation flips. Heavier usage doesn&#8217;t mean a heavier invoice. Once the box is yours, running more agents, longer reasoning, bigger context, all of it lives inside a cost you already know. The incentive inverts: instead of being penalized for using AI more, you&#8217;re free to. That&#8217;s the difference between intelligence as a metered utility and intelligence as owned capability.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">A few things follow from owning the stack instead of renting it:</p>
<p class="font-claude-response-body break-words whitespace-normal leading-[1.7]"><strong>Your costs are knowable before the work starts.</strong> A flat monthly fee means the budget conversation happens once, up front, not in a panicked review when the usage report comes in. No surprises, no variable cloud bill, no quarter blown in a month.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-[1.7]"><strong>Your usage can succeed without punishing you.</strong> The whole point of AI is to do more with it over time. On a metered model, success is the thing that breaks your budget. On owned infrastructure, success is just success.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-[1.7]"><strong>You&#8217;re not locked to one vendor&#8217;s pricing whims.</strong> Microsoft just moved a product to token billing because its own costs ran away. When you don&#8217;t own the layer your business depends on, someone else&#8217;s cost problem becomes your pricing problem overnight. We run the model that fits the job, on hardware that&#8217;s ours, so a vendor&#8217;s repricing isn&#8217;t your emergency.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-[1.7]"><strong>Your data stays yours.</strong> This was always the foundation. Models run locally, on our infrastructure, in our facility. Your data never routes through someone else&#8217;s cloud to get answered. Your data, your rules. The cost predictability is a benefit that rides on top of the same architecture that keeps your information private in the first place.</p>
<h2 class="text-text-100 mt-3 -mb-1 text-[1.125rem] font-bold">The meter was always the business model</h2>
<p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">The token-billing crisis isn&#8217;t a bug in cloud AI. It&#8217;s the business model working as designed. Usage was always going to climb, agents were always going to multiply the consumption, and the bill was always going to follow the curve. May just happened to be the month some very large companies looked up and noticed.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">The organizations that come out of this ahead won&#8217;t be the ones who used AI the least to survive the bill. They&#8217;ll be the ones who stopped renting intelligence by the token and started owning it, so that using more was never the thing that hurt them.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">If you&#8217;re staring at an AI bill that grows every time the tools actually work, that&#8217;s worth a conversation. We&#8217;re always happy to compare notes on what fixed-cost, private AI looks like for an organization your size. You can reach us at <a class="underline underline underline-offset-2 decoration-1 decoration-current/40 hover:decoration-current focus:decoration-current" href="https://modtechgroup.com/consultation/">modtechgroup.com/consultation</a>.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">Because when the token bill finally comes due across the industry, you want to be the company that already knows its number.</p>
<hr class="border-border-200 border-t-0.5 my-3 mx-1.5" />
<p class="font-claude-response-body break-words whitespace-normal leading-[1.7]"><em>Modular Technology Group builds and operates private AI infrastructure on owned, US-based hardware: fixed pricing, local inference, your data and your AI under your rules, from dirt to desktop. <a class="underline underline underline-offset-2 decoration-1 decoration-current/40 hover:decoration-current focus:decoration-current" href="https://modtechgroup.com/">modtechgroup.com</a></em></p>
<p>The post <a href="https://modtechgroup.com/when-the-token-bill-comes-due-what-uber-and-microsoft-just-taught-the-rest-of-us-about-renting-intelligence/">When the Token Bill Comes Due: What Uber and Microsoft Just Taught the Rest of Us About Renting Intelligence</a> appeared first on <a href="https://modtechgroup.com">Modular Technology Group</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>When Google Validates Your Architecture: Private AI Was Never the Alternative</title>
		<link>https://modtechgroup.com/when-google-validates-your-architecture-private-ai-was-never-the-alternative/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=when-google-validates-your-architecture-private-ai-was-never-the-alternative</link>
		
		<dc:creator><![CDATA[Arthur]]></dc:creator>
		<pubDate>Mon, 27 Apr 2026 18:06:32 +0000</pubDate>
				<category><![CDATA[Private AI]]></category>
		<category><![CDATA[hyperscalers]]></category>
		<category><![CDATA[own vs rent]]></category>
		<guid isPermaLink="false">https://modtechgroup.com/?p=5765</guid>

					<description><![CDATA[<p>At Google Cloud Next 2026 in Las Vegas this week, Google made a quiet but significant announcement: Gemini can now run on a single air-gapped server, fully disconnected from the internet — and from Google itself.</p>
<p>The post <a href="https://modtechgroup.com/when-google-validates-your-architecture-private-ai-was-never-the-alternative/">When Google Validates Your Architecture: Private AI Was Never the Alternative</a> appeared first on <a href="https://modtechgroup.com">Modular Technology Group</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p><div class="fusion-fullwidth fullwidth-box fusion-builder-row-1 fusion-flex-container nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-padding-top:40px;--awb-padding-bottom:40px;--awb-flex-wrap:wrap;" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1310.4px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"><div class="fusion-layout-column fusion_builder_column fusion-builder-column-0 fusion_builder_column_1_1 1_1 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:100%;--awb-margin-top-large:0px;--awb-spacing-right-large:1.92%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:1.92%;--awb-width-medium:100%;--awb-order-medium:0;--awb-spacing-right-medium:1.92%;--awb-spacing-left-medium:1.92%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-1"><figure class="wp-block-image size-large modular-cdn-hero"><img decoding="async" src="https://assets.modtechgroup.com/blog/concept-art/2026/04/pitch-0744-20260427-135632.png" alt="When Google Validates Your Architecture: Private AI Was Never the Alternative" /></figure>
<p>At Google Cloud Next 2026 in Las Vegas this week, Google made a quiet but significant announcement: Gemini can now run on a single air-gapped server, fully disconnected from the internet — and from Google itself.</p>
<p>The product is a Dell-certified, Google-approved hardware appliance delivered through a neocloud partner called Cirrascale Cloud Services. Eight Nvidia GPUs. Confidential computing protections. The marketing hook: &#8220;pull the plug and the model vanishes.&#8221;</p>
<p>We&#8217;ve been watching the coverage with genuine interest. And a fair bit of déjà vu.</p>
<h2>The Market Just Caught Up</h2>
<p>For years, enterprise organizations in financial services, healthcare, defense, and government faced what analysts called an impossible tradeoff: access the most powerful AI models through public cloud APIs — and surrender control of your data — or settle for less capable open-source models you could host yourself.</p>
<p>Google&#8217;s announcement is a formal acknowledgment that this framing was always wrong. The demand for fully private AI wasn&#8217;t a niche concern. It was the only architecturally honest answer for any organization that takes data governance seriously.</p>
<p>Modular Technology Group has been building on that premise since before it was a keynote slide.</p>
<h2>What Google Is Actually Selling</h2>
<p>Let&#8217;s be precise about the offering, because the details matter.</p>
<p>The Cirrascale deployment requires a Google-certified hardware platform. It requires a partnership with a specific neocloud provider. It requires Google&#8217;s approval of the appliance configuration. General availability is projected for June or July 2026 — it&#8217;s in preview now.</p>
<p>And the selling point — that the model &#8220;vanishes when you pull the plug&#8221; — is a confidential computing feature that ties the model weights to the specific hardware. Impressive engineering. But consider what it implies: you are still dependent on Google&#8217;s certification ecosystem to acquire and maintain access to the model. The sovereignty is physical, not architectural.</p>
<p>The right question for any enterprise evaluating this: <strong>What is your exit strategy?</strong></p>
<ul>
<li>What happens if Cirrascale changes its pricing or partnership terms?</li>
<li>What happens if Google deprecates the on-premises licensing tier?</li>
<li>What happens when the certified hardware goes end-of-life?</li>
</ul>
<p>Vendor lock-in doesn&#8217;t disappear because the server is in your rack. It moves from the network layer to the hardware and licensing layer.</p>
<h2>A Different Architectural Bet</h2>
<p>Modular Technology Group made a different set of choices when we designed our private AI infrastructure.</p>
<p><strong>Model-agnostic.</strong> We are not tied to any single model provider. Our clients run the models that fit their use case — whether that&#8217;s an open-weight model, a fine-tuned variant, or a frontier model accessed under controlled conditions. When a better model ships, you switch. No re-certification. No new appliance.</p>
<p><strong>Hardware-agnostic.</strong> We operate in a FedRAMP-authorized data center on infrastructure you control. You are not locked to a specific GPU configuration or a vendor-approved hardware stack. The architecture scales with your needs, not with a product roadmap you don&#8217;t control.</p>
<p><strong>Fixed, transparent pricing.</strong> No usage-based API billing. No surprise invoices at the end of the month. You know what you&#8217;re paying. That predictability is a feature, not an accident.</p>
<p><strong>Available now.</strong> Not in preview. Not GA in Q3. Running, deployed, with clients in production today.</p>
<h2>Data Sovereignty Is Architecture, Not Proximity</h2>
<p>The broader lesson from Google&#8217;s announcement isn&#8217;t about Google. It&#8217;s about how the enterprise AI market is maturing in its understanding of what &#8220;private&#8221; actually means.</p>
<p>Physical proximity — a server in your building, or in a data center you can point to — is necessary but not sufficient. True data sovereignty requires architectural ownership: control over the model, the infrastructure, the data pipeline, and the exit path.</p>
<p>When your AI model &#8220;vanishes when you pull the plug,&#8221; ask yourself: whose plug is it, really?</p>
<p>At Modular Technology Group, &#8220;Your Data, Your Rules&#8221; isn&#8217;t a product announcement. It&#8217;s been the design constraint from the beginning.</p>
<p>If you&#8217;re evaluating private AI infrastructure — whether in response to this week&#8217;s news or because you&#8217;ve been thinking about it longer than Google has been announcing it — we&#8217;re happy to compare architectures.</p>
<p><a href="https://modtechgroup.com/consultation/">Schedule a conversation →</a></p>
<hr />
<p><em>Source inspiration: <a href="https://venturebeat.com/technology/googles-gemini-can-now-run-on-a-single-air-gapped-server-and-vanish-when-you-pull-the-plug" target="_blank" rel="noopener">LinkedIn</a></em></p>
</div></div></div></div></div></p>
<p>The post <a href="https://modtechgroup.com/when-google-validates-your-architecture-private-ai-was-never-the-alternative/">When Google Validates Your Architecture: Private AI Was Never the Alternative</a> appeared first on <a href="https://modtechgroup.com">Modular Technology Group</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>The Model That Barely Slows Down: Gemma 4 26B vs Qwen 3.6 35B at Long Context</title>
		<link>https://modtechgroup.com/the-model-that-barely-slows-down-gemma-4-26b-vs-qwen-3-6-35b-at-long-context/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=the-model-that-barely-slows-down-gemma-4-26b-vs-qwen-3-6-35b-at-long-context</link>
		
		<dc:creator><![CDATA[Arthur]]></dc:creator>
		<pubDate>Wed, 22 Apr 2026 16:40:03 +0000</pubDate>
				<category><![CDATA[AI Infrastructure]]></category>
		<category><![CDATA[benchmarks]]></category>
		<category><![CDATA[models]]></category>
		<category><![CDATA[self hosting]]></category>
		<guid isPermaLink="false">https://modtechgroup.com/?p=5743</guid>

					<description><![CDATA[<p>We ran Gemma 4 26B and Qwen 3.6 35B-A3B head-to-head on the same server, same quantization, same protocol. Gemma 4 is 3.7× faster at 32k context — and 7.2× faster at 128k. The gap widens with context, and the reason reveals something important about model selection for long-context workloads.</p>
<p>The post <a href="https://modtechgroup.com/the-model-that-barely-slows-down-gemma-4-26b-vs-qwen-3-6-35b-at-long-context/">The Model That Barely Slows Down: Gemma 4 26B vs Qwen 3.6 35B at Long Context</a> appeared first on <a href="https://modtechgroup.com">Modular Technology Group</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p><div class="fusion-fullwidth fullwidth-box fusion-builder-row-2 fusion-flex-container nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-padding-top:40px;--awb-padding-bottom:40px;--awb-flex-wrap:wrap;" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1310.4px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"><div class="fusion-layout-column fusion_builder_column fusion-builder-column-1 fusion_builder_column_1_1 1_1 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:100%;--awb-margin-top-large:0px;--awb-spacing-right-large:1.92%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:1.92%;--awb-width-medium:100%;--awb-order-medium:0;--awb-spacing-right-medium:1.92%;--awb-spacing-left-medium:1.92%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-2"><h2>The Model That Barely Slows Down: Gemma 4 26B vs Qwen 3.6 35B at Long Context</h2>
<p><em>Modular Technology Group · April 22, 2026</em></p>
<hr />
<p>I&#8217;ve been thinking a lot about what it means to deploy a model in production versus benchmark it in a controlled setting. Most benchmarks pick short prompts — 1k, 2k tokens — and declare a winner. That&#8217;s fine for answering quick questions. It&#8217;s irrelevant if you&#8217;re building anything real: document analysis, long-thread summarization, multi-turn reasoning agents, whole-repo code review.</p>
<p>So we don&#8217;t benchmark that way.</p>
<p>Two days ago we published numbers for Qwen 3.6 35B-A3B across 32k, 64k, and 128k contexts on our dedicated AI server. Today we ran the same protocol against Google&#8217;s brand-new <strong>Gemma 4 26B</strong> — same hardware, same quantization, same prompts, same three-context sweep.</p>
<p>The headline: <strong>Gemma 4 26B is 3.7× faster than Qwen 3.6 35B-A3B at 32k context. At 128k, it&#8217;s 7.2× faster.</strong></p>
<p>And unlike Qwen 3.6 — which we watched degrade from 26 tokens/sec at 32k to 9 tokens/sec at 128k — Gemma 4 barely moves. It went from 96 to 87 to 65 tokens/sec. The curve is nearly flat. That changes the infrastructure calculus entirely.</p>
<hr />
<h2>The Setup</h2>
<p>Same hardware as Monday&#8217;s Qwen bench. Same server. Same protocol.</p>
<table class="wp-block-table is-style-stripes">
<thead>
<tr>
<th>Platform</th>
<th>Hardware</th>
<th>Engine</th>
<th>Quantization</th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>&#8220;Reach&#8221;</strong> — dedicated AI server</td>
<td>2× NVIDIA RTX 4070 Ti, 24 GB VRAM total</td>
<td>Ollama (llama.cpp/GGUF)</td>
<td>Q4_K_M</td>
</tr>
</tbody>
</table>
<p>Models under test:</p>
<ul>
<li><strong>Gemma 4 26B</strong> (Google, MoE A4B — ~4B active parameters per token, 26B total)</li>
<li><strong>Qwen 3.6 35B-A3B</strong> (Alibaba, MoE A3B — ~3B active parameters per token, 36B total)</li>
</ul>
<p>Protocol matches <a href="../2026-04-20-qwen36-forge-vs-reach/">our April 20 baseline bench</a>:</p>
<ul>
<li>Context windows: <strong>32k, 64k, 128k tokens</strong></li>
<li>Prompt: synthetic filler at <strong>85% of target context budget</strong> — same bytes both models</li>
<li>Completion: <strong>256 tokens</strong>, temperature 0.1</li>
<li>Trials: <strong>3 measured per cell</strong> (+ 1 warm-up discarded per model×context)</li>
<li>Model unloaded between runs — no contamination from the other model&#8217;s KV cache</li>
<li>Explicit <code>num_ctx</code> override on every Ollama request (Ollama silently caps at 4,096 without it — we learned this the hard way and documented it)</li>
</ul>
<p>18 total runs. 0 failures. Variance: under 1% across all trials.</p>
<hr />
<h2>The Numbers</h2>
<table class="wp-block-table is-style-stripes">
<thead>
<tr>
<th>Context</th>
<th>Gemma 4 26B</th>
<th>Qwen 3.6 35B-A3B</th>
<th>Gemma advantage</th>
</tr>
</thead>
<tbody>
<tr>
<td>32k</td>
<td><strong>96.4 tok/s</strong> · 3.5s wall</td>
<td>26.3 tok/s · 11.5s wall</td>
<td><strong>3.7×</strong></td>
</tr>
<tr>
<td>64k</td>
<td><strong>86.7 tok/s</strong> · 4.1s wall</td>
<td>19.3 tok/s · 15.7s wall</td>
<td><strong>4.5×</strong></td>
</tr>
<tr>
<td>128k</td>
<td><strong>65.2 tok/s</strong> · 5.9s wall</td>
<td>9.1 tok/s · 31.4s wall</td>
<td><strong>7.2×</strong></td>
</tr>
</tbody>
</table>
<p>Let me frame the wall-clock numbers concretely. A 256-token response — roughly one dense paragraph — takes:</p>
<ul>
<li>Gemma 4 at <strong>any</strong> context: under 6 seconds</li>
<li>Qwen 3.6 at 32k: 11.5 seconds</li>
<li>Qwen 3.6 at 64k: 15.7 seconds</li>
<li>Qwen 3.6 at 128k: <strong>31.4 seconds</strong></li>
</ul>
<p>That&#8217;s the difference between a tool you hold a conversation with and one you fire off while you pour another cup of coffee.</p>
<p><img decoding="async" class="aligncenter size-full" src="https://modtechgroup.com/wp-content/uploads/2026/04/chart-1-gen-throughput-1.png" alt="Generation throughput by context — Gemma 4 vs Qwen 3.6" /></p>
<hr />
<h2>The Surprise Finding: The Architecture Gap Widens With Context</h2>
<p>Here&#8217;s where I want to spend more time, because this isn&#8217;t just a &#8220;new model is faster&#8221; story.</p>
<p>Both models are Mixture-of-Experts. Both use Q4_K_M quantization. Both run on the same two GPUs. At 32k context, the gap is already 3.7×. By 128k, it&#8217;s 7.2×. The gap nearly doubles as the context grows.</p>
<p>Why?</p>
<p><img decoding="async" class="aligncenter size-full" src="https://modtechgroup.com/wp-content/uploads/2026/04/chart-3-degradation-1.png" alt="Throughput degradation curve — how each model handles growing context" /></p>
<p><strong>Qwen 3.6 35B-A3B:</strong></p>
<ul>
<li>36B total parameters, ~3B active per token</li>
<li>At 128k context, generation drops to 9.1 tok/s</li>
<li>Degradation from 32k to 128k: <strong>-65%</strong></li>
</ul>
<p><strong>Gemma 4 26B:</strong></p>
<ul>
<li>26B total parameters, ~4B active per token</li>
<li>At 128k context, generation holds at 65.2 tok/s</li>
<li>Degradation from 32k to 128k: <strong>-32%</strong></li>
</ul>
<p>The KV cache grows linearly with context. At 128k, both models are operating under the same VRAM pressure we documented Monday — memory bandwidth is the bottleneck, not compute. The GPUs are reading enormous amounts of data per generated token.</p>
<p>The difference is the underlying architecture. Gemma 4&#8217;s A4B configuration activates more parameters per token than Qwen 3.6&#8217;s A3B, which would normally suggest higher compute overhead. But the total parameter count is smaller (26B vs 36B), meaning the weight tensors being loaded from VRAM on each generation step are physically smaller. Less data to move per token. Less memory bandwidth consumed per token. The gap widens with context precisely because the bandwidth-bound regime amplifies parameter-count differences.</p>
<p>In short: <strong>at long context, smaller total parameter count beats higher active parameter count</strong> when you&#8217;re memory-bandwidth constrained.</p>
<p>This is the kind of finding that doesn&#8217;t show up in a 2k-token benchmark.</p>
<p><img decoding="async" class="aligncenter size-full" src="https://modtechgroup.com/wp-content/uploads/2026/04/chart-5-speedup.png" alt="Gemma 4 speedup multiplier grows with context" /></p>
<hr />
<h2>What This Means for Infrastructure Selection</h2>
<p>The previous bench taught us that hardware tier matters: dual mid-range GPUs on a dedicated server outperformed an M4 Max laptop by 5.3× at 128k. This bench teaches something different — that <strong>model architecture matters just as much as hardware</strong> for long-context workloads.</p>
<p>A few things I&#8217;m taking away from this:</p>
<p><strong>Context length changes the whole model-selection calculus.</strong> Qwen 3.6 35B-A3B is an excellent model. For reasoning tasks at moderate contexts, it&#8217;s still compelling. But if your workload involves 64k+ prompts — and an increasing number of real workloads do — the throughput differential is severe enough to matter operationally. A 7.2× speed penalty at 128k context isn&#8217;t a marginal difference; it&#8217;s a different class of tool.</p>
<p><strong>Model architecture is an infrastructure decision, not just a capability decision.</strong> When selecting a model for a production deployment, we now explicitly consider the active-parameter count, total parameter count, and their ratio alongside benchmark capability scores. Two MoE models with similar benchmark performance can behave completely differently under sustained long-context load.</p>
<p><strong>The bandwidth-bottleneck pattern generalizes.</strong> We saw last week that at 128k context, the GPUs were running at 6–7% compute utilization with VRAM saturated at 91%. The compute was idle. The memory bus was the choke. Gemma 4 takes advantage of this constraint by keeping its weight tensor smaller — it&#8217;s effectively doing less memory I/O per token, which is exactly what you want when the memory bus is your ceiling.</p>
<p><strong>Smaller isn&#8217;t always slower.</strong> The conventional wisdom is that a 36B model is &#8220;better&#8221; than a 26B model — more parameters, more capacity. For generation throughput under memory-bandwidth constraints, the relationship inverts. Whether Gemma 4 produces better *output quality* than Qwen 3.6 for a given task is a separate question — one worth benchmarking rigorously — but on pure throughput at long context, the smaller model wins decisively.</p>
<hr />
<h2>An Honest Note on Prompt Eval Telemetry</h2>
<p>In our Qwen benchmark, Ollama reported prompt ingestion speeds of 20k–44k tokens/sec across the three context sizes — a useful data point for pipeline latency estimation.</p>
<p>For Gemma 4, Ollama&#8217;s <code>prompt_eval_duration</code> consistently reported 13–19ms across all three context windows, implying millions of tokens/sec. This is a KV-cache reuse artifact: the warm-up trial primes the cache, and subsequent trials appear to skip most or all of the ingestion phase. We&#8217;re reporting this honestly rather than publishing the inflated numbers. The wall-clock timing captures the full end-to-end latency accurately; the prompt_eval field in Ollama&#8217;s response for Gemma 4 requires more investigation before we&#8217;d cite it confidently.</p>
<p>What we can say: if Gemma 4 is achieving genuine KV-cache reuse across sequential requests with the same prompt prefix, that&#8217;s actually a meaningful throughput advantage for multi-turn workloads. We&#8217;ll dig into this in a follow-up run with cold-cache isolation.</p>
<hr />
<h2>What&#8217;s Next</h2>
<p>Two tests I want to run before I&#8217;m satisfied this benchmark is complete:</p>
<p>1. <strong>Cold-cache prompt eval isolation for Gemma 4</strong> — force model reload between every trial to get a clean first-ingestion measurement 2. <strong>Output quality comparison</strong> — throughput advantage is only relevant if the output quality holds up. We&#8217;ll run a structured evaluation comparing Gemma 4 and Qwen 3.6 on legal document analysis and long-form synthesis tasks — the actual workloads our clients care about</p>
<p>The throughput finding is real and significant. Whether Gemma 4 earns its place in the production stack depends on the quality side of the equation.</p>
<hr />
<h2>The Infrastructure View</h2>
<p>For organizations evaluating private AI: the model landscape is moving fast, and the performance characteristics of new models don&#8217;t always fit the pattern of what came before. A model selection decision from six months ago might be suboptimal today — not because the old model got worse, but because the new options are sufficiently different architecturally.</p>
<p>This is part of why we run these benchmarks with our own hardware and real workloads rather than relying on published leaderboard numbers. Leaderboards optimize for benchmark performance. We care about throughput under the memory constraints of actual production hardware, at the context lengths real workloads require.</p>
<p>Modular doesn&#8217;t resell AI. We build, host, and run the infrastructure ourselves — which means we&#8217;re measuring what actually matters to us operationally. These numbers are real because they have to be.</p>
<p>If you&#8217;re working through a private AI infrastructure decision and want to compare notes, I&#8217;m always open to the conversation.</p>
<hr />
<h2>Appendix: Methodology &amp; Caveats</h2>
<p><strong>Models:</strong></p>
<ul>
<li>Gemma 4 26B (Google DeepMind) — Ollama tag <code>gemma4:26b</code>, GGUF Q4_K_M, 17.9 GB, digest <code>5571076f3d70</code></li>
<li>Qwen 3.6 35B-A3B (Alibaba) — Ollama tag <code>qwen3.6:35b-a3b</code>, GGUF Q4_K_M, 23.9 GB, digest <code>07d35212591f</code></li>
</ul>
<p><strong>Hardware:</strong> ReachAI server, 2× NVIDIA RTX 4070 Ti (12 GB VRAM each, 24 GB total), Ubuntu 24.04, Ollama v0.11.10</p>
<p><strong>Prompt construction:</strong> Same filler text (140-char repeating unit), calibrated to 85% of target token budget. Tokenizer calibration ran 2026-04-22: Gemma 4 measures 6.76 chars/token, Qwen 3.6 measures 6.81 chars/token — within 1%. Same prompt bytes sent to both models; both reported nearly identical <code>prompt_eval_count</code> (26,639 vs 26,632 at 32k; 53,259 vs 53,252 at 64k; 106,479 vs 106,472 at 128k), confirming tokenizer parity.</p>
<p><strong>Trial structure:</strong> 1 warm-up trial discarded per model×context cell (model remained loaded for context-cache consistency), then 3 measured trials. Model explicitly unloaded between models using Ollama&#8217;s <code>keep_alive: 0</code> mechanism to prevent cross-contamination.</p>
<p><strong>Completion:</strong> 256 tokens, <code>temperature=0.1</code>.</p>
<p><strong>Ollama:</strong> Explicit <code>num_ctx</code> override per request. Default silently caps at 4,096 tokens.</p>
<p><strong>Variance:</strong> Under 1% across all trials for both models. Gemma 4: 96.4 / 96.6 / 96.4 at 32k; 65.2 / 65.3 / 65.2 at 128k. Rock solid.</p>
<p><strong>Caveats:</strong></p>
<ul>
<li>Gemma 4 prompt eval timing is not reported due to KV-cache reuse masking cold-start latency in Ollama. Wall-clock timing is accurate.</li>
<li>Both models use Q4_K_M GGUF — same quantization scheme, though the underlying weight distributions differ.</li>
<li>Tests executed on a dedicated server with no competing workloads.</li>
<li>Output quality comparison not included in this benchmark — throughput only.</li>
</ul>
<p><strong>Reproducibility:</strong> All scripts and raw data archived at <code>reports/bench-archive/2026-04-22-gemma4-vs-qwen36/</code>. Benchmark harness is argparse-driven and can be re-run against any Ollama endpoint.</p>
<hr />
<p><em>Cale Hollingsworth is the founder of Modular Technology Group, which builds and hosts private AI workspaces in a FedRAMP data center. He has been advising organizations on infrastructure strategy since 1993. LinkedIn: Fractional CTO | Private AI Infrastructure Strategist | Evangelist @ Modular Technology Group | RAG Architect | Future-Proofing Organizations Since 1993</em></p>
<p><em>#PrivateAI #DataPrivacy #yourdatayourrules</em></p>
</div></div></div></div></div></p>
<p>The post <a href="https://modtechgroup.com/the-model-that-barely-slows-down-gemma-4-26b-vs-qwen-3-6-35b-at-long-context/">The Model That Barely Slows Down: Gemma 4 26B vs Qwen 3.6 35B at Long Context</a> appeared first on <a href="https://modtechgroup.com">Modular Technology Group</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Same AI Model, Two Hardware Tiers — And Why Context Length Is the Hidden Variable</title>
		<link>https://modtechgroup.com/same-ai-model-two-hardware-tiers-and-why-context-length-is-the-hidden-variable/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=same-ai-model-two-hardware-tiers-and-why-context-length-is-the-hidden-variable</link>
		
		<dc:creator><![CDATA[Arthur]]></dc:creator>
		<pubDate>Tue, 21 Apr 2026 01:59:58 +0000</pubDate>
				<category><![CDATA[AI Infrastructure]]></category>
		<category><![CDATA[benchmarks]]></category>
		<category><![CDATA[context length]]></category>
		<category><![CDATA[hardware]]></category>
		<guid isPermaLink="false">https://modtechgroup.com/?p=5732</guid>

					<description><![CDATA[<p>We put Qwen 3.6 35B-A3B on a developer laptop and a dual-GPU server. The speed gap grows from 2.4× to 5.3× as context grows — and the real bottleneck turns out not to be compute.</p>
<p>The post <a href="https://modtechgroup.com/same-ai-model-two-hardware-tiers-and-why-context-length-is-the-hidden-variable/">Same AI Model, Two Hardware Tiers — And Why Context Length Is the Hidden Variable</a> appeared first on <a href="https://modtechgroup.com">Modular Technology Group</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p><div class="fusion-fullwidth fullwidth-box fusion-builder-row-3 fusion-flex-container nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-padding-top:40px;--awb-padding-bottom:40px;--awb-flex-wrap:wrap;" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1310.4px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"><div class="fusion-layout-column fusion_builder_column fusion-builder-column-2 fusion_builder_column_1_1 1_1 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:100%;--awb-margin-top-large:0px;--awb-spacing-right-large:1.92%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:1.92%;--awb-width-medium:100%;--awb-order-medium:0;--awb-spacing-right-medium:1.92%;--awb-spacing-left-medium:1.92%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-3"><h2>Same AI Model, Two Hardware Tiers — And Why Context Length Is the Hidden Variable</h2>
<p><em>Modular Technology Group · April 20, 2026</em></p>
<hr />
<p>Ask any AI vendor how fast their stack runs and you&#8217;ll get a single headline number. &#8220;40 tokens per second.&#8221; &#8220;Under a second to first token.&#8221; Impressive — until you realize the benchmark prompt was 200 words long and you&#8217;re planning to feed it a 300-page document.</p>
<p>This week we took <strong>Qwen 3.6 35B-A3B</strong> — a state-of-the-art Mixture-of-Experts model released a few days ago — and pointed it at two very different pieces of hardware. Same model. Same questions. Same quantization tier (4-bit). Only the hardware changed.</p>
<p>The result isn&#8217;t just a horse race. It&#8217;s a quiet lesson in why the specs that matter for AI aren&#8217;t always the specs that get advertised.</p>
<hr />
<h2>Why We Ran This</h2>
<p>At Modular, we route the same model across different infrastructure depending on the workload. A developer laptop handles quick, short-context tasks. A dedicated AI server handles long-document analysis, multi-turn agent reasoning, and anything that needs a big context window.</p>
<p>The question isn&#8217;t &#8220;which is faster.&#8221; A server beats a laptop. That&#8217;s boring.</p>
<p>The real question: <strong>at what context length does routing to the dedicated server become worth it?</strong> Without numbers, every routing decision is a guess. So we measured.</p>
<hr />
<h2>The Setup</h2>
<table class="wp-block-table is-style-stripes">
<thead>
<tr>
<th>Platform</th>
<th>Hardware</th>
<th>Engine</th>
<th>Quantization</th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>&#8220;Forge&#8221;</strong> — developer laptop</td>
<td>MacBook Pro M4 Max, 64 GB unified memory</td>
<td>LM Studio (MLX backend)</td>
<td>MLX 4-bit</td>
</tr>
<tr>
<td><strong>&#8220;Reach&#8221;</strong> — dedicated AI server</td>
<td>2× NVIDIA RTX 4070 Ti, 24 GB VRAM total</td>
<td>Ollama v0.11.10 (llama.cpp/GGUF)</td>
<td>Q4_K_M GGUF</td>
</tr>
</tbody>
</table>
<p>We ran the model at three context sizes — <strong>32k, 64k, and 128k tokens</strong> — and measured how long each host took to generate a 256-token response. Three trials per cell. Temperature fixed at 0.1 for near-determinism. Prompt content matched byte-for-byte. Tokenizer output cross-checked. Apples to apples.</p>
<p>For the 128k results, we added six total trials across two independent sessions to nail the number down.</p>
<hr />
<h2>Result #1: One Host Stays Usable. The Other Doesn&#8217;t.</h2>
<p><img decoding="async" class="aligncenter size-full" src="https://modtechgroup.com/wp-content/uploads/2026/04/chart-3-degradation.png" alt="Throughput degradation as context grows — Forge collapses, Reach holds up" /></p>
<p>At 32k context, both platforms deliver workable performance. The MacBook runs at 10.8 tokens/sec — slower than the dedicated server, but perfectly fine for interactive chat.</p>
<p>Then the context grows.</p>
<p>At 64k, the MacBook drops to <strong>4.8 tokens/sec</strong>. At 128k, it collapses to <strong>1.7 tokens/sec</strong>.</p>
<p>The dedicated server, meanwhile, holds its shape:</p>
<table class="wp-block-table is-style-stripes">
<thead>
<tr>
<th>Context</th>
<th>Forge (MBP)</th>
<th>Reach (Dual GPU)</th>
<th>Reach advantage</th>
</tr>
</thead>
<tbody>
<tr>
<td>32k</td>
<td>10.8 tok/s</td>
<td>26.3 tok/s</td>
<td><strong>2.4× faster</strong></td>
</tr>
<tr>
<td>64k</td>
<td>4.8 tok/s</td>
<td>19.3 tok/s</td>
<td><strong>4.0× faster</strong></td>
</tr>
<tr>
<td>128k</td>
<td>1.7 tok/s</td>
<td>9.0 tok/s</td>
<td><strong>5.3× faster</strong></td>
</tr>
</tbody>
</table>
<p>Notice the pattern: the gap widens with every doubling of context. This isn&#8217;t a flat advantage — it compounds. By the time you&#8217;re at 128k, the kind of window you need for whole-document analysis or agent reasoning, the server is over five times faster than the laptop.</p>
<p><img decoding="async" class="aligncenter size-full" src="https://modtechgroup.com/wp-content/uploads/2026/04/chart-1-gen-throughput.png" alt="Generation throughput at 32k, 64k, and 128k across both hosts" /></p>
<hr />
<h2>Result #2: The Honest Metric Is Wall Time</h2>
<p>Tokens-per-second is abstract. What does this actually feel like to a human waiting for an answer?</p>
<p><img decoding="async" class="aligncenter size-full" src="https://modtechgroup.com/wp-content/uploads/2026/04/chart-2-wall-time.png" alt="Wall-clock time for a 256-token response" /></p>
<p>A <strong>256-token reply</strong> — roughly one solid paragraph — takes:</p>
<table class="wp-block-table is-style-stripes">
<thead>
<tr>
<th>Context</th>
<th>Forge</th>
<th>Reach</th>
</tr>
</thead>
<tbody>
<tr>
<td>32k</td>
<td>23 seconds</td>
<td><strong>12 seconds</strong></td>
</tr>
<tr>
<td>64k</td>
<td>54 seconds</td>
<td><strong>15 seconds</strong></td>
</tr>
<tr>
<td>128k</td>
<td><strong>2 minutes, 32 seconds</strong></td>
<td><strong>31 seconds</strong></td>
</tr>
</tbody>
</table>
<p>That&#8217;s the difference between a tool you can hold a conversation with and a tool you fire off and check back on later.</p>
<hr />
<h2>Result #3: The Bottleneck Isn&#8217;t What You Think</h2>
<p>Here&#8217;s where it gets interesting.</p>
<p>During the 128k runs on the dedicated server, we monitored both GPUs continuously. The VRAM was pegged — <strong>22.2 GB of 24 GB total, 91% saturation</strong>. So the GPUs must have been pegged too, right?</p>
<p>Not even close.</p>
<p><img decoding="async" class="aligncenter size-full" src="https://modtechgroup.com/wp-content/uploads/2026/04/chart-5-gpu-util.png" alt="GPU compute utilization vs VRAM saturation during 128k inference" /></p>
<p>The two GPUs, theoretically capable of hundreds of trillions of operations per second, sat at <strong>6–7% utilization</strong>. They weren&#8217;t waiting for work. They were waiting for <em>memory</em>.</p>
<p>At long context lengths, the model has to read the entire &#8220;KV cache&#8221; — every token it&#8217;s seen so far — to generate each new token. Enormous quantities of data move between VRAM and the compute cores every few milliseconds. The memory bus becomes the choke point long before the math does.</p>
<p>This is the single most important finding in the entire exercise, because it reframes how to evaluate future hardware.</p>
<p><strong>More FLOPS won&#8217;t fix this.</strong> When the question becomes &#8220;should we buy the next card when it drops?&#8221; — the answer starts with its memory bandwidth spec, not its TFLOPS number. That&#8217;s the opposite of what most marketing collateral emphasizes.</p>
<hr />
<h3>The Same Story, Live From Production Telemetry</h3>
<p><img decoding="async" class="aligncenter size-full" src="https://modtechgroup.com/wp-content/uploads/2026/04/chart-6-grafana-live.jpg" alt="Live Grafana capture during the 128k verification runs" /></p>
<p>This is real production monitoring from our own dashboard during the benchmark — not synthetic charts. Three things worth noticing:</p>
<ul>
<li><strong>Both GPU panels are nearly identical.</strong> Both cards track the same 5–7% load pattern. That&#8217;s tensor parallelism working.</li>
<li><strong>The staircase in &#8220;Total Memory Used.&#8221;</strong> Each step is a single 128k trial committing its KV cache, then holding it. Three trials, three plateaus, climbing toward the 24 GB ceiling.</li>
<li><strong>Compute is flat. Memory is climbing.</strong> The shape of the real data tells the same story as the synthetic chart: this workload lives and dies by memory, not by compute.</li>
</ul>
<p>This is the visibility that separates production AI infrastructure from &#8220;we installed it and hope it works.&#8221;</p>
<hr />
<h2>Result #4: Tensor Parallelism Done Right</h2>
<p>One thing the dedicated server does exceptionally well: split the model cleanly across both GPUs.</p>
<p><img decoding="async" class="aligncenter size-full" src="https://modtechgroup.com/wp-content/uploads/2026/04/chart-4-vram-split.png" alt="Per-GPU VRAM split at 128k — textbook balance" /></p>
<p>At 128k context, the memory load is nearly identical on both cards — <strong>11,101 MiB on GPU 0, 11,117 MiB on GPU 1</strong>. A difference of 16 MiB out of over 11,000. That&#8217;s Ollama&#8217;s tensor-parallel splitter working exactly as designed. No card is bearing extra load. No GPU is OOMing. No spillover to CPU.</p>
<p>Tensor parallelism isn&#8217;t automatic. It requires compatible hardware, deliberate configuration, and a runtime that actually supports it. It&#8217;s also invisible to the end user — which is exactly how it should be.</p>
<hr />
<h2>What This Means for How You Deploy AI</h2>
<p>If you&#8217;re prototyping against 4k-to-16k prompts on a decent laptop, you&#8217;re fine. For a team running real AI applications against real-world documents, the math shifts quickly.</p>
<p>A few honest observations from this data:</p>
<ul>
<li><strong>Context length matters more than model size.</strong> A 35B-parameter model can feel snappy or geological depending entirely on how much context you feed it. Marketing benchmarks rarely mention this.</li>
<li><strong>Hardware choice is a memory problem, not just a compute problem.</strong> Two mid-range GPUs with balanced VRAM can outperform much more expensive single-GPU setups for long-context work.</li>
<li><strong>Consumer hardware has real limits.</strong> M-series Macs are remarkable for the price. But physics is physics. There&#8217;s a reason production AI workloads live on dedicated servers.</li>
<li><strong>Private infrastructure isn&#8217;t only about sovereignty.</strong> It&#8217;s also about having the right hardware for the right context, predictable performance, and the ability to scale without a surprise cloud bill.</li>
</ul>
<p>At Modular, we deploy private AI infrastructure that gets these details right — matching the model, the quantization, the hardware, and the runtime so answers come back in seconds, not minutes. Data stays private. Costs stay fixed. Performance stays predictable.</p>
<p>Your data, your rules. Your hardware, matched to your workload.</p>
<hr />
<h2>Appendix: Methodology &amp; Caveats</h2>
<p><strong>Model:</strong> Qwen 3.6 35B-A3B (Mixture-of-Experts — 36B total parameters, 3B active per token)</p>
<p><strong>Prompts:</strong> Synthetic filler text sized to 85% of target context, with a single consistent question appended. Byte-identical across both hosts. Tokenizer output verified to match (<code>prompt_tokens</code> reported identically on each side).</p>
<p><strong>Trials:</strong> Three per context-size × host cell for the primary run. Six additional trials at 128k on the dedicated server across two independent sessions. Variance across all six 128k runs: under 2% (8.94–9.03 tok/s).</p>
<p><strong>Completion target:</strong> 256 tokens, <code>temperature=0.1</code>.</p>
<p><strong>Ollama configuration:</strong> Explicit <code>num_ctx</code> override on every request. Default caps context at 4,096 tokens — enough to silently invalidate every long-context test if you miss it.</p>
<p><strong>Caveats:</strong></p>
<ul>
<li>Quantization formats differ (MLX 4-bit vs Q4_K_M GGUF). Both are 4-bit but not bit-identical.</li>
<li>The MacBook was running normal background workloads during the test, not dedicated. A clean bench would improve its numbers modestly but not flip the conclusion.</li>
<li>Single model tested. Different architectures — dense transformers, larger MoEs, specialized coding models — will scale differently.</li>
<li>The 6–7% GPU utilization figure reflects generation phase only. Prompt evaluation phase utilization was much higher, but brief.</li>
</ul>
<p><strong>Raw data and all benchmark scripts:</strong> Available on request. Fully reproducible.</p>
<hr />
<p><em>Modular Technology Group builds and hosts private AI workspaces with open-source components, in a FedRAMP data center. We use what we sell.</em></p>
</div></div></div></div></div></p>
<p>The post <a href="https://modtechgroup.com/same-ai-model-two-hardware-tiers-and-why-context-length-is-the-hidden-variable/">Same AI Model, Two Hardware Tiers — And Why Context Length Is the Hidden Variable</a> appeared first on <a href="https://modtechgroup.com">Modular Technology Group</a>.</p>
]]></content:encoded>
					
		
		
			</item>
	</channel>
</rss>
