Infrastructure

The stack behind it

Everything here runs on hardware we operate. No managed platform, no per-seat SaaS holding the content hostage, and no figure on this page typed by hand — the numbers are counted off the machines and stamped with the date they were counted.

Measured 2026-08-30.

Compute

Dedicated servers across two providers, plus a workstation node that carries part of the ingest load.

10
servers and a VPS
8
orchestration instances
490
running containers
4.0 TB
provisioned storage
1.5 TB
in use
354
container images

Data

Every database is self-hosted. Nothing sits in a vendor's console waiting for a pricing change.

193
database instances
178
PostgreSQL
11
Redis / Valkey
4
other engines
20
object-storage buckets
199
persistent volumes

PostGIS and pgvector where the work needs geometry or embeddings. Backups run nightly to local disk and to object storage — two destinations, because one backup is not a backup — with 14-day local and 60-day offsite retention.

Automation

The estate maintains itself. Jobs harvest, verify, re-measure and alert without anyone opening a terminal.

766
scheduled jobs
170
orchestrated tasks
596
system cron entries
139
deploy webhooks
139
managed applications
147
Git repositories
4.4K
deployments shipped

Every application deploys from Git on push — 4,391 of them since 2026-06, which is when the current orchestration was stood up rather than when the work began. The first repository dates to 2024-03. Health, freshness and screenshot checks run nightly and only speak up when something changes state.

Artificial intelligence

Models we host ourselves, on our own GPUs, at no per-token cost — alongside a platform that fronts hundreds more.

5
local model families
4
inference servers
220 GB
model weights on disk
400+
models via vincony.com
5.1M
entities in the knowledge graph
45M
pages in the knowledge graph

Generation, embeddings and reranking all run locally on llama.cpp. vincony.com puts 400+ models from every major lab behind one subscription, and its knowledge graph is built from Wikipedia, Wikidata and 27 public data sources.

Acquisition

Four browser engines and three egress paths, combined into nine escalating strategies. It starts free and only spends money when a page genuinely resists.

4
browser engines
3
egress paths
9
combined strategies
500
dedicated proxy IPs
5K GB
proxy bandwidth / month
217
liveness-tested APIs

The engines are a self-hosted crawl API, a headless-browser crawler, an undetected-Chrome solver and a stealth Firefox build. Each can egress from residential Thailand, from European datacentres, or through geo-targeted residential proxies — and it is the ENGINE × EGRESS combinations that form the ladder, not nine separate tools. A commercial browser sits at the end as an opt-in last resort, and in practice is almost never reached. Every external API in the catalogue was tested for liveness rather than trusted from a published list.

Self-hosted services

Roughly two dozen third-party services run on our own machines rather than someone else's.

24
third-party services
3
crawl API instances
18
mail-server containers
250
mailboxes being provisioned
3
monitoring systems
1
programmatic video renderer

Mail runs on Mailcow — Postfix, Dovecot, Rspamd, SOGo, ClamAV and Unbound, on our own hardware, with 250 mailboxes currently being provisioned onto it. Alongside it: self-hosted error tracking, uptime monitoring, deployment orchestration, a search-and-vector database layer, a graph database, an object store, a document renderer, a change-detection watcher, a geocoder, a translation engine, an identity server, a forum, and a programmatic video renderer. Transactional mail goes out through a dedicated provider across 99 verified sending domains.

Outreach and data harvesting

Systematic acquisition, not a bought list. A category-by-city grid swept until it stopped returning anything new.

5.2K
verified business contacts
1.9K
grid queries executed
37
business categories
9
concurrent map-scraper instances
51M
pages across the estate
492
screenshots on file

Every contact was harvested from a public listing and then verified — the run stopped when a full enrichment pass over the remaining candidates returned zero new results, which is the honest definition of finished rather than an arbitrary target. Sending is throttled and consent-aware.

Observability

One panel watches every site, every deploy and every error across the estate.

58
uptime monitors
101
error-tracked projects
399
DNS zones
2.6K
DNS records
1.1K
hostnames
14
languages served

A single control panel aggregates uptime, error tracking, deploy state and SEO signals for the whole estate, so a failure anywhere surfaces in one place — one screen rather than ten. It tracks every application across every host, flags a deploy that silently rolled back, and surfaces errors captured by the self-hosted error tracker. 16 properties publish in more than one language, measured from what they actually serve rather than from what a field said, and mail leaves through 132 verified sending domains.

Search and content operations

A semi-automated SEO system reads real search-performance data for every property, identifies where a page is close to breaking through, and produces a prescription an engineer can apply directly. It is deliberately semi-automatic: the system finds and proposes, a human approves, and the change is then verified against the same data it came from.

Content pipelines run the same way. Nightly jobs measure what each property publishes, record when it last changed, and check that what it serves is what it claims — but they are barred from rewriting a word of copy on their own, because a job that silently rewrites content is how a library fills up with text nobody can source.

Compute, data, automation and observability figures measured 2026-08-30. Proxy pool, API catalogue and workstation figures reviewed 2026-08-10. Sites in the collection are counted live.