We moved three live models to new GPUs without stopping them
- Where it started
- No external APIs, and the models in service could not go down for even a minute.
- What we did
- Rebuilt on-prem serving and moved each model from H100 to DGX in turn.
Shipped in banking, public sector and SaaS
No external APIs allowed, a mountain of HWP files, a token bill that grows every month. We wire stalled AI into your internal systems and keep running it until your people use it every day.
Where our team has done the work (client names stay confidential)
01 · Services
It won’t connect to your systems. Security says no. Costs spike. Reviewers want proof. We take each one on separately, with fixed scope and deliverables.
We connect Slack, groupware and databases over MCP, and log who approved what.
We set up GPU servers and serve models on them. We have even moved live models to new hardware without stopping service.
We extract tables and paragraphs without breaking them, attach the source to every answer, then cross-check it.
Leaner prompts, caching, small models for easy requests, and fixes for slow queries.
We build eval datasets, score them automatically and prepare the documents for TTA technical review.
We look at your workflows, data and security rules, and come back with one task that pays off and a fixed quote.
02 · Cases
Work our team designed, built and ran ourselves. We can’t name clients, but every figure was measured in production.
03 · Products
We use them every day, so they were battle-tested first. See everything at products.ferrorium.com.
A macOS editor that puts nine agents in a 3×3 grid and gives them all work at once.
Request →04 · Process
Prove it small, then expand. Before each step, we agree on how success will be measured.
We look at workflows, data and security rules and pick the one task that clearly pays off.
We build a version that runs on your real data and score it against the agreed bar.
We add permissions, logging and monitoring, then deploy to cloud or an air-gapped network.
We watch cost and quality, then hand it over so your team can run it.
05 · FAQ
Last updated
We take AI that stalled after the demo and put it to work. That means internal AI agents, air-gapped on-prem LLMs, HWP document AI and LLM cost cuts, and we stay through operations instead of stopping at design. We also build and use our own developer tools: ctx, FoldChat and Nine Agents.
Yes. We run open models on your own GPU servers, so data has nowhere to go. Inside a commercial bank’s segregated network, we moved three live models from H100 to DGX without a single minute of downtime.
Yes. We extract HWP and HWPX with tables and paragraphs intact and use them for search and verification. On a public AI evaluation project we pulled out 100% of the HWP spec, word for word, and passed TTA technical review.
Usually, yes. We trim prompts, add caching, route easy requests to smaller models and fix slow queries. For a sales email SaaS, that cut serving cost 38%, output tokens 49% and p95 latency 70%.
Usually one week to assess and two to three weeks to pilot, so you see something working within four weeks. Cost depends on scope and environment; after the assessment you get a fixed quote that won’t creep.
Use the form at the bottom of this page. Pick a request type and write two or three lines about your workflow and data setup. We reply by email within two business days. If the form fails, email [email protected] directly.
Two or three lines is enough. Within two business days we email back what we can take on and a quote.
A good fit if