HomeAICantonese AI model built for $250,000 to reach 80 million speakers

Cantonese AI model built for $250,000 to reach 80 million speakers

Ask any major AI model a question in English or Mandarin, and it will likely answer with fluency and nuance. Ask it something in Cantonese, and the cracks start to show. That gap is exactly what a Hong Kong startup called Votee AI is trying to close, building a Cantonese AI model designed to serve banks, universities, and government departments in a language the biggest AI labs have largely overlooked.

Key takeaways

  • Most leading AI models, built by companies like OpenAI, Anthropic, DeepSeek, Moonshot AI and Z.ai, are developed primarily in English and Mandarin because their makers are based in the US or mainland China.
  • Cantonese is spoken by more than 80 million people worldwide, yet it has a much smaller pool of standardized digital text than Mandarin, making it a textbook “low resource language” for AI training.
  • Votee AI retrains open-weight models such as Meta’s Llama on Cantonese data, growing its training corpus from 100 million to more than 500 million tokens through scraping, community sources and synthetic data.
  • Its resulting models run around 70 billion parameters, smaller than frontier systems, and cost roughly $250,000 to train using between 500 million and 1 billion tokens, a fraction of what English-language models require.
  • The company frames its work as part of a broader “sovereign AI” push, and is now in talks to expand into Southeast Asia through AI Singapore.

AI’s English and Mandarin Dominance Leaves Cantonese Behind

The current AI boom runs almost entirely on two languages. According to Pak-Sun Ting, CEO of Votee AI, “the whole AI revolution is in English and Mandarin,” and “there’s only a very small fraction that represents other languages.” That imbalance isn’t accidental. The companies leading the field — OpenAI, Anthropic, DeepSeek, Moonshot AI, Z.ai and others — are almost all headquartered in the United States or mainland China, so it’s no surprise their flagship models are strongest in their creators’ native tongues.

That leaves dozens of widely spoken languages, and even more regional dialects, trailing far behind in AI capability. Cantonese is a striking example of just how large that gap can be, even for a language spoken by tens of millions of people every day.

Why Cantonese Speaks a Different Language Than Mandarin

Cantonese is often described as a “dialect” of Chinese, but that label understates how distinct it really is. It has different grammar from Mandarin and its own vocabulary, and in Hong Kong, speakers frequently blend English and Cantonese words within the same sentence. More than 80 million people speak Cantonese globally, a figure roughly on par with the number of Korean speakers, and larger than the populations speaking Italian or Thai.

Despite that scale, Cantonese hasn’t produced anywhere near the volume of standardized written text that Mandarin or English have. Benchmarks such as HKCanto-Eval, developed by researchers at Kyushu University and the Education University of Hong Kong alongside the local AI community hon9kon9ize, and sponsored by Votee, found that mainstream AI models can handle everyday Cantonese reasonably well but routinely stumble on cultural and local knowledge. As Ting puts it, Cantonese “is used in education, healthcare, and police communications,” so when AI fails to cover it properly, “AI is essentially useless.”

Inside Votee AI’s Approach to Building a Cantonese AI Model

Rather than building a model from the ground up, Votee takes an existing open-weight system, such as Meta’s Llama or Alibaba’s Qwen, and retrains it heavily on Cantonese-language data. Ting describes the process as “essentially taking the same steps as if you were training a model from scratch,” just built on top of someone else’s foundation. The retrained systems are then sold to banks, universities and government departments that need AI tools capable of working in Cantonese rather than defaulting to Mandarin or English.

Sourcing Data From Broadcasters and Synthetic Sets

Building a usable corpus for a low resource language AI project means scraping together data from wherever it can be found. Votee pulls Cantonese text through online scraping, including content from Radio Television Hong Kong, the city’s public broadcaster. It supplements that with material from universities, the wider community, and archives from its earlier work as a big data company. Where real-world text runs short, Votee generates its own synthetic Cantonese data sets to fill the gaps. Combined, those efforts pushed the company’s Cantonese corpus from 100 million tokens up to more than 500 million tokens.

Model Size and Training Costs Compared to Frontier Labs

Votee’s resulting models sit at around 70 billion parameters, noticeably smaller than the leading systems from top labs. Even so, Ting says the models are capable enough to understand and reason in Cantonese for the practical tasks its clients need. The economics tell their own story: Votee trains its models on between 500 million and 1 billion tokens, at an estimated cost of roughly $250,000. Frontier English-language models, by contrast, are trained on trillions of tokens at vastly higher expense. That price gap is exactly why a smaller, targeted Cantonese AI model can be commercially viable even for a startup rather than a tech giant.

This cost structure matters beyond Votee itself. It suggests that serving a low resource language AI market doesn’t require the billions poured into frontier labs — it requires a workable open-weight base model, a solid data pipeline, and a fraction of the compute. That’s a meaningful signal for other underserved languages watching how Cantonese AI development plays out.

Sovereign AI and the Push Into Southeast Asia

Votee’s work sits inside a wider trend known as sovereign AI: the idea that governments and companies want to own their own data, models and infrastructure rather than depending entirely on providers based overseas. “AI has become such an essential need, and so you don’t want to be tethered to anybody else who can turn it off,” Ting says.

Owning the Stack, Even Partially

Ting is upfront that the fullest version of sovereign AI, where a country controls every link of the AI supply chain, is “very difficult” to achieve in practice. His more realistic suggestion is that the emphasis of nations is on possessing both the underlying foundation models and the software solutions constructed from them. He also notes that governments rarely need frontier-level power to automate specific tasks — even a model with as few as 1 billion parameters can suit narrower government use cases, or a smaller local-language layer can simply route outputs from a larger English or Chinese model.

Hardware Partnerships and What’s Next

Votee doesn’t limit itself to a single ecosystem. Ting says the company works with models from MiniMax and SenseTime and can run operations on Nvidia chips. “We can use Nvidia chips, we can use Moonshot or DeepSeek’s model,” he says. “We’re that person in high school who’s friends with everyone.” Votee describes itself as profitable in the sense that its revenues exceed costs, funded largely through client contracts, and counts Hong Kong tycoon Allan Zeman, known for developing the city’s Lan Kwai Fong nightlife district, as an advisor.

The company’s ambitions now stretch beyond Hong Kong. Ting says Votee is in active discussions with AI Singapore and plans to expand further across Southeast Asia. He’s also eyeing a longer-term goal: using AI to help protect endangered languages in regions including East Asia, North America and Africa. Ting frames the stakes in stark terms, calling the dominance of English-language AI a “typewriter moment,” where productivity gains are so large that people abandon their own language just to keep up. “People will adopt English just because the typewriter’s productivity is so strong versus their own language,” he says. Whether a 70-billion-parameter Cantonese AI model can meaningfully push back against that pull remains an open question, but for Ting, the motivation runs deeper than market share. “Every language that dies, you lose another way of seeing the world,” he says.

FAQ

Why is Votee AI focusing on Cantonese for AI model development?

Because Cantonese, spoken by over 80 million people, is significantly different from Mandarin and underserved in AI, which affects sectors like education and healthcare where accurate local-language tools are essential.

How does Votee AI build its Cantonese AI models?

Votee AI retrains open-weight models from developers like Meta using Cantonese data collected through scraping, community sources, universities, and synthetic data, growing its corpus from 100 million to over 500 million tokens.

What is sovereign AI and how does Votee AI relate to it?

Sovereign AI refers to governments and companies owning their AI data, models, and infrastructure rather than relying entirely on foreign providers. Votee AI supports this idea by localizing AI models for Cantonese and partnering with local and global hardware and model providers.

What are Votee AI’s plans for regional expansion?

Votee AI is in active discussions with AI Singapore and plans to expand into Southeast Asia, while also exploring how AI could help protect endangered languages in other parts of the world.

Article produced with the assistance of artificial intelligence and reviewed by the editorial team.

Francesco Antonio Russo
Web 3.0 entrepreneur for over 4 years, expert in Cryptocurrencies and Artificial Intelligence. He uses his cross-functional skills for functional and trend-following Social Media Management.
RELATED ARTICLES

Stay updated on all the news about cryptocurrencies and the entire world of blockchain.

Featured video

LATEST