# Ferrorium Last updated: 2026-10-04 > Ferrorium is a Korean AI engineering team that takes AI stalled after the demo and puts it to work, all the way through operations. We build internal AI agents, air-gapped on-premise LLMs, document AI for HWP/HWPX with cited RAG, and LLM cost cuts. Our team has done this work for a commercial bank, public agencies, a public research institute, an overseas securities firm and B2B SaaS. Client names stay confidential. ## What we take on Most AI pilots don't fail on the model. They stall because they won't connect to internal systems, security rules out external APIs, costs spike, or reviewers want proof. We handle each of these separately: - AI agents: agents you call from Slack, connected to groupware and databases over MCP, with approval and audit logs. Our internal platform runs 143 MCP tools and is used by 15 of 21 staff. - On-prem and air-gapped LLMs: GPU servers (H100, DGX), model serving, and hardware moves without downtime. - Document AI and RAG: HWP/HWPX extraction with tables and paragraphs intact (100% word-level spec coverage on a public project), answers with sources, cross-checking. - LLM cost and performance: leaner prompts, caching, model routing and query tuning. On one project: p95 latency −70%, serving cost −38%, output tokens −49%. - AI evaluation and certification: eval datasets, automated scoring, TTA technical review documents. ## What happened on real projects (measured in production) - Commercial bank: we moved three live models from H100 to DGX inside a segregated network with zero minutes of downtime. - Public AI evaluation support system: manual evaluation became an automated pipeline; HWP specs extracted 100% word for word; passed TTA technical review. - Sales email SaaS: p95 latency −70%, serving cost −38%, output tokens −49%; 2.2 million emails sent and 1,701 signups. - Overseas securities firm: financial data queries went from 411ms to 1.6ms. - Internal AI agent platform: 143 MCP tools, 639 PRs opened by agents, used by 15 of 21 staff. - Supply-chain security scan: from 568 seconds to 0.7 seconds. ## How an engagement runs One week to assess, two to three weeks to pilot, so you see something working within four weeks. Then 4–10 weeks to build, followed by operations and handover to your team. After the assessment you get a fixed quote. ## Our own tools - [ctx](https://ctx.ferrorium.com/): a free CLI coding agents use to find and read files, keeping only the lines that matter. - [FoldChat](https://foldchat.ferrorium.com/): an Android app to keep driving Claude Code on your own computer from your phone. - [Nine Agents](https://nineagents.ferrorium.com/): a macOS editor that runs nine agents in a 3×3 grid at once. ## Pages - [Home (Korean)](https://ferrorium.com/) - [Home (English)](https://ferrorium.com/en/) ## Contact Use the form at https://ferrorium.com/#contact (English: https://ferrorium.com/en/#contact). Pick a request type (one of the five services, a one-week assessment, ctx for teams, FoldChat waitlist, Nine Agents beta, or other) and write two or three lines about your situation. We reply by email within two business days. If the form fails: wks0968@gmail.com