Shipped in banking, public sector and SaaS

The demo got applause.
Nobody uses it yet.

No external APIs allowed, a mountain of HWP files, a token bill that grows every month. We wire stalled AI into your internal systems and keep running it until your people use it every day.

Where our team has done the work (client names stay confidential)

  • Commercial bankOn-prem LLM
  • Public sectorAI evaluation · B2G
  • Public researchDocument AI
  • Overseas securitiesPerformance
  • B2B SaaSAI automation
  • Internal AIAgent platform

01 · Services

The model was never the problem.
These five things were.

It won’t connect to your systems. Security says no. Costs spike. Reviewers want proof. We take each one on separately, with fixed scope and deliverables.

  • AI agent development

    Agents you can call from Slack

    We connect Slack, groupware and databases over MCP, and log who approved what.

    15/21staff who actually use the internal platform
    Request →
  • On-prem · air-gapped LLM

    LLMs that never send a byte outside

    We set up GPU servers and serve models on them. We have even moved live models to new hardware without stopping service.

    0 mindowntime moving 3 models H100→DGX
    Request →
  • Document AI · RAG

    Search that reads HWP and cites its sources

    We extract tables and paragraphs without breaking them, attach the source to every answer, then cross-check it.

    100%word-level HWP spec coverage
    Request →
  • LLM cost · performance

    Same quality, smaller bill

    Leaner prompts, caching, small models for easy requests, and fixes for slow queries.

    −38%LLM serving cost
    Request →
  • AI evaluation · certification

    Answer “does it really work?” with numbers

    We build eval datasets, score them automatically and prepare the documents for TTA technical review.

    TTAtechnical review passed
    Request →
  • Not sure where to start?

    Give us one week

    We look at your workflows, data and security rules, and come back with one task that pays off and a fixed quote.

02 · Cases

Not slide-deck numbers.
Production numbers.

Work our team designed, built and ran ourselves. We can’t name clients, but every figure was measured in production.

FinanceCommercial bank

We moved three live models to new GPUs without stopping them

Where it started
No external APIs, and the models in service could not go down for even a minute.
What we did
Rebuilt on-prem serving and moved each model from H100 to DGX in turn.
0 mindowntime moving 3 models from H100 to DGX
Ask about a similar project →
PublicAI evaluation support system

Manual AI evaluation became automatic, and passed TTA review

Where it started
Evaluation was done by hand, and the evidence was buried deep in HWP specs.
What we did
Extracted HWP specs word by word and fed them into an eval pipeline with cross-checks.
100%word-level HWP spec extraction · TTA technical review passed
Ask about a similar project →
SaaSSales email automation

Faster emails, then 2.2 million of them went out

Where it started
Generation was slow, and token costs grew with every send.
What we did
Redesigned prompts, caching and model routing, and tuned slow queries.
−70%p95 latency · serving cost −38% · output tokens −49% · 1,701 signups
Ask about a similar project →
  • Overseas securities firmFinancial data queries went from 411ms to 1.6ms
  • Internal AI agent platform143 MCP tools, 639 PRs opened by agents, used by 15 of 21 staff
  • Supply-chain security scanA 568-second scan now takes 0.7 seconds

03 · Products

We kept rebuilding these tools.
So we shipped them.

We use them every day, so they were battle-tested first. See everything at products.ferrorium.com.

ctx ● Free · public

A CLI coding agents use to find and read files. It keeps only the lines that matter, saving context and tokens.

Request →

FoldChat Waitlist

Step away from your desk and keep driving Claude Code on your computer from your Android phone.

Request →

Nine Agents Beta soon

A macOS editor that puts nine agents in a 3×3 grid and gives them all work at once.

Request →

04 · Process

Something working in four weeks

Prove it small, then expand. Before each step, we agree on how success will be measured.

  1. 1 week

    Assess

    We look at workflows, data and security rules and pick the one task that clearly pays off.

  2. 2–3 weeks

    Pilot

    We build a version that runs on your real data and score it against the agreed bar.

  3. 4–10 weeks

    Build

    We add permissions, logging and monitoring, then deploy to cloud or an air-gapped network.

  4. Ongoing

    Run · hand over

    We watch cost and quality, then hand it over so your team can run it.

05 · FAQ

What people ask before they start

Last updated

What does Ferrorium actually do?

We take AI that stalled after the demo and put it to work. That means internal AI agents, air-gapped on-prem LLMs, HWP document AI and LLM cost cuts, and we stay through operations instead of stopping at design. We also build and use our own developer tools: ctx, FoldChat and Nine Agents.

Can it run on a network with no internet?

Yes. We run open models on your own GPU servers, so data has nowhere to go. Inside a commercial bank’s segregated network, we moved three live models from H100 to DGX without a single minute of downtime.

Can you read HWP files?

Yes. We extract HWP and HWPX with tables and paragraphs intact and use them for search and verification. On a public AI evaluation project we pulled out 100% of the HWP spec, word for word, and passed TTA technical review.

Can you cut the LLM bill we already have?

Usually, yes. We trim prompts, add caching, route easy requests to smaller models and fix slow queries. For a sales email SaaS, that cut serving cost 38%, output tokens 49% and p95 latency 70%.

How long does it take, and what does it cost?

Usually one week to assess and two to three weeks to pilot, so you see something working within four weeks. Cost depends on scope and environment; after the assessment you get a fixed quote that won’t creep.

How do we get in touch?

Use the form at the bottom of this page. Pick a request type and write two or three lines about your workflow and data setup. We reply by email within two business days. If the form fails, email [email protected] directly.

Just tell us where you’re stuck

Two or three lines is enough. Within two business days we email back what we can take on and a quote.

A good fit if

  • Your pilot is done but nobody uses it
  • Security rules out external APIs
  • The answers are buried in HWP or PDF files
  • Your LLM bill keeps growing
Install ctx free

We use this information only to reply to your inquiry.