Open-source inference at production scale, served entirely from within Southeast Asia — dedicated endpoints, transparent per-token pricing, and zero vendor lock-in for teams building in Singapore, Indonesia, Thailand, Vietnam, Malaysia and the Philippines.
Deploy Llama, Qwen, DeepSeek and GPT-OSS on dedicated endpoints with sub-second targets and 99.95% uptime. Autoscaling, speculative decoding and routing across our Southeast Asia data centers hold latency steady from your first request to your hundred-billionth.
Explore endpoints →Move from a single prototype to hundreds of millions of tokens a minute without re-architecting. Autoscaling and routing across our Southeast Asia endpoints hold latency steady the whole way.
Transparent per-token rates across shared and dedicated tiers, with volume discounts that apply automatically as usage grows. No idle GPU charges.
Text, code, vision and reasoning models behind one API. Mix modalities in the same pipeline without juggling providers or contracts.
Native function calling, structured JSON output and built-in safety guardrails, so agents behave predictably once they're talking to the real world.
Adapt a base model to your data, then deploy the checkpoint straight to a Hi Liberty endpoint with the same per-token pricing and SLAs as everything else.
High-performance embedding models and vector storage sit next to inference, so indexing, retrieval and generation stay inside one governed platform.
Text, code and vision models behind a single OpenAI-compatible API, backed by production SLAs and predictable per-token pricing.
Learn more →Explore inference history, curate datasets from production traffic, and move straight into your next fine-tuning pass.
Learn more →Streamlined fine-tuning workflows that take an open model from your data to a live, billable endpoint in one pass.
Learn more →Independently benchmarked throughput and latency across flagship models, so procurement doesn't have to take our word for it.
A 99.95% uptime SLA, autoscaling and speculative decoding keep response times flat under real production load, not just in demos.
LLMs, vision, reasoning and embedding models under a single account and a single invoice — expanding every month.
Point your existing client at Hi Liberty and switch models with a string, not a migration.
Read the docs →from openai import OpenAI client = OpenAI( base_url="https://api.hiliberty.ai/v1", api_key=HI_LIBERTY_API_KEY ) completion = client.chat.completions.create( model="deepseek-v3.2", messages=[{ "role": "user", "content": "What does liberty mean for infrastructure?" }] ) print(completion.choices[0].message.content) # -> "Owning the stack you build on."
Start on shared access, move to dedicated endpoints with reserved capacity and a 99.95% SLA whenever you need it. Transparent $/token throughout, no idle infrastructure costs.
Yes. Hi Liberty is built for large-scale, production-grade AI workloads. Dedicated endpoints deliver sub-second inference, a 99.95% uptime SLA, and autoscaling throughput for workloads well past a hundred million tokens a minute.
Scale from experimentation to full production across our Southeast Asia footprint, without rate throttles or GPU management of your own.
We onboard new open-source releases regularly, based on customer demand and independent benchmarking. Enterprise users can request model optimization or custom deployment support directly through our Solutions team.
Yes. Dedicated endpoints provide guaranteed isolation, predictable latency and reserved compute capacity, including a 99.95% SLA, custom autoscaling and a choice of deployment location within our Singapore, Jakarta or Bangkok facilities. Contact our team to size an instance for your workload and compliance needs.
Deploy directly through the dashboard or API. Each deployment runs with transparent per-token pricing and inherits the same security and latency guarantees as standard endpoints. Full post-training workflows let you go from dataset to deployed checkpoint in one pass.
A zero-retention mode is available, where requests and outputs are never stored or reused for training. Data is processed exclusively within our Southeast Asia facilities (Singapore, Jakarta and Bangkok), certified to SOC 2 Type II and ISO 27001, and never leaves the region.
The free tier ships with generous defaults; enterprise plans remove caps entirely. We size endpoints to your traffic profile — reach out and we'll scope an endpoint for your real-world workload.
Hi Liberty currently serves customers based in Southeast Asia only — including Singapore, Indonesia, Malaysia, Thailand, Vietnam and the Philippines. All infrastructure, billing and support are local to the region. We don't yet support sign-ups or billing outside Southeast Asia; reach out if you'd like to be notified when that changes.
Pricing is transparent per-token, with clear input/output separation and volume discounts as you scale. No hidden infrastructure or idle GPU costs — you pay only for what you serve.