The global token factory
TAPEX is an AI cloud built for one job: generating tokens at scale. Latest-generation NVIDIA compute, high-throughput inference, and cost-per-token economics — served from Thailand to the world.
Verified customers · Aligned with US export controls (EAR/BIS)
Built on the frontier of accelerated computing
TAPEX deploys NVIDIA's latest-generation accelerated computing as its foundational compute layer — engineered for inference performance per watt and per dollar.
NVIDIA GB300 Grace Blackwell at the core
Our clusters pair Grace CPUs with Blackwell-generation GPUs on a coherent memory fabric — purpose-matched to large-model inference and fine-tuning, and surrounded by a multi-tier GPU fleet so every workload runs on right-sized silicon.
Blackwell-generation performance
Early access to GB300-class hardware puts TAPEX at the frontier of tokens per second — with RTX PRO 6000 and GeForce RTX 5090 tiers for workloads that don't need rack scale.
Serving-optimised stack
Continuous batching, paged KV-cache and per-model quantisation tuned across the fleet — throughput without quality loss.
Performance per watt
Efficient power and cooling design translates directly into lower, more predictable cost per token.
A factory measured in tokens, not racks
We run our infrastructure the way a factory runs a production line — utilisation, yield and unit cost are the metrics that matter.
Optimised cost per token
Every layer of the stack — hardware selection, batching, scheduling, power — is tuned toward one number: what a million tokens costs to produce.
High utilisation
Dense scheduling and workload mixing keep GPUs producing instead of idling — utilisation is our core discipline.
Efficient power & cooling
Facility design tuned for Blackwell-class density, so energy goes into tokens — not overhead.
Throughput, guaranteed
Capacity reservations with clear throughput commitments — plan token volumes like any other supply contract.
Transparent supply
A trusted, verified counterparty in the NVIDIA channel — customers know exactly where their compute comes from.
Compute, the way you need it
Inference-as-a-service
High-throughput serving for production LLM workloads, priced per token.
- OpenAI-compatible endpoints
- Batch & real-time
- Per-token pricing
GPU-as-a-service
Dedicated Blackwell-class GPU capacity, bare-metal control when you need it.
- Reserved instances
- Full-node allocation
- Custom images
Fine-tuning
Adapt open and custom models on large-memory systems built for long context and large batches.
- LoRA & full fine-tunes
- Private datasets stay in-region
Dedicated capacity
Long-term reserved clusters for model providers and enterprises with sustained token demand.
- Committed throughput
- Predictable unit cost
How a token factory works
A factory turns raw inputs into finished goods. Ours turns energy, silicon and model weights into tokens — the unit every AI product is built from.
Raw inputs
Power, cooling and Blackwell-class silicon — provisioned, dense and ready. The plant floor of the factory.
Models loaded
Open-weight and customer models resident in GPU memory, quantised and compiled for the hardware they run on.
Orchestration
Continuous batching and smart scheduling mix workloads to keep every GPU producing — utilisation is yield.
Tokens delivered
Streamed through low-latency APIs, metered per token — with throughput you can plan against like any supply contract.
Energy+Silicon+Models=Tokens at scale
The open-weight frontier, on tap
Production-grade serving for the model families teams actually deploy — or bring your own weights and we'll run them.
Llama
MetaThe default open-weight family for chat, agents and enterprise assistants.
DeepSeek
DeepSeekFrontier-class reasoning at open-weight economics — built for hard problems.
Qwen
AlibabaStrong multilingual coverage — including Thai and Southeast Asian languages.
Mistral
Mistral AIEfficient dense and MoE models with excellent cost-to-quality ratios.
Gemma
GoogleLightweight models that shine on high-volume, latency-sensitive workloads.
GPT-OSS
OpenAIOpen-weight reasoning models with broad ecosystem and tooling support.
Bring your own model
Custom checkpoints, fine-tunes and proprietary architectures — deployed on dedicated capacity, weights and data kept in-region.
Your inference layer, in your region
Hyperscalers serve the world from somewhere else. TAPEX serves it from Thailand — with a home-ground advantage for Southeast Asia: a regional alternative to Western neoclouds, built for the workloads and regulations of this market.
Lower latency for ASEAN users
In-region serving cuts round-trips to distant clouds — critical for real-time products.
Data residency
Workloads and datasets stay in Thailand, simplifying regional data-governance requirements.
Built for regional scale
Capacity for AI startups, enterprises and model providers across Thailand and ASEAN.
A counterparty you can verify
Advanced compute carries responsibility. TAPEX operates with full transparency toward NVIDIA channel partners, regulators and customers.
Customer verification
Every customer completes KYC before access. We know who runs on our infrastructure — always.
End-use verification
Workloads are screened against declared end uses, with ongoing monitoring for diversion risk.
US export-control alignment
Full alignment with EAR/BIS regulations — a trusted, transparent counterparty for NVIDIA channel partners and enterprise customers.
Start producing tokens at scale
Tell us about your workload — model, token volumes, latency needs — and our team will come back with capacity and pricing.