Tencent Hy Research Team Debuts Hy4 Preview: Flagship 770B-A49B Architecture Built for High-Density Agentic WorkflowsThe Tencent Hy research team has officially released the preview version of its next-generation foundation model, Hy4. Markedly shifting strategy from the smaller, budget-focused design of its predecessor, Hy3, Tencent’s new release enters the frontier class delivering benchmark performance on par with leading Chinese AI flagships including DeepSeek V4 Pro, Kimi K3, GLM-5.3, and Qwen3.8 Max.
Built on a massive 770B-A49B Mixture-of-Experts (MoE) architecture, Hy4 dramatically expands parameter capacity while incorporating a native 1-Million Token Context Window. This extended memory capacity allows the model to maintain context across long-horizon reasoning tasks and handle complex multi-step technical execution without losing track of instructions.
Despite the significant increase in parameter scale, Tencent has maintained a strong price-to-performance advantage:
Standard API Rates: Priced at $0.834 per million input tokens and $2.501 per million output tokens (reflecting a multi-fold increase over Hy3 due to its frontier parameter footprint).
Ultra-Low Prompt Caching: Set at just $0.042 per million tokens, representing only a minor bump over Hy3.
This low prompt-caching cost drastically lowers operational overhead for agentic workflows, where agents repeatedly re-read system prompts, environment logs, and long tool-use histories.
Tencent's strategic shift in model development, while previous Hy versions aimed for lightweight and low-cost API deployments, demonstrates Tencent's intention to compete directly at the highest level of underlying models. The switch to 770 billion total parameters enables the network to store knowledge of the broader world and handle sophisticated reasoning tasks that smaller models typically struggle with.
Autonomous AI agents operate through iterative loops, passing system instructions, function trees, and past procedures back and forth in each iteration. In traditional pricing models, repeatedly reading large context windows on every call inflates costs. By reducing alert caching fees to $0.042 per million tokens, Tencent allows developers to run complex, multi-iterion agent pipelines at a fraction of the standard API cost.
Although Hy4 boasts 770 billion total parameters, Mixture-of-Experts (MoE) routing enables only 49 billion parameters per token. This selective enablement strategy keeps inference latency manageable. And by balancing GPU memory bandwidth requirements, tokens can be generated quickly even when processing long and contextually complex alert messages.
Source: Tencent Hy
Tencent Hy Research Team Debuts Hy4 Preview: Flagship 770B-A49B Architecture Built for High-Density Agentic WorkflowsThe Tencent Hy research team has officially released the preview version of its next-generation foundation model, Hy4. Markedly shifting strategy from the smaller, budget-focused design of its predecessor, Hy3, Tencent’s new release enters the frontier class delivering benchmark performance on par with leading Chinese AI flagships including DeepSeek V4 Pro, Kimi K3, GLM-5.3, and Qwen3.8 Max.
Built on a massive 770B-A49B Mixture-of-Experts (MoE) architecture, Hy4 dramatically expands parameter capacity while incorporating a native 1-Million Token Context Window. This extended memory capacity allows the model to maintain context across long-horizon reasoning tasks and handle complex multi-step technical execution without losing track of instructions.
Despite the significant increase in parameter scale, Tencent has maintained a strong price-to-performance advantage:
Standard API Rates: Priced at $0.834 per million input tokens and $2.501 per million output tokens (reflecting a multi-fold increase over Hy3 due to its frontier parameter footprint).
Ultra-Low Prompt Caching: Set at just $0.042 per million tokens, representing only a minor bump over Hy3.
This low prompt-caching cost drastically lowers operational overhead for agentic workflows, where agents repeatedly re-read system prompts, environment logs, and long tool-use histories.
Tencent's strategic shift in model development, while previous Hy versions aimed for lightweight and low-cost API deployments, demonstrates Tencent's intention to compete directly at the highest level of underlying models. The switch to 770 billion total parameters enables the network to store knowledge of the broader world and handle sophisticated reasoning tasks that smaller models typically struggle with.
Autonomous AI agents operate through iterative loops, passing system instructions, function trees, and past procedures back and forth in each iteration. In traditional pricing models, repeatedly reading large context windows on every call inflates costs. By reducing alert caching fees to $0.042 per million tokens, Tencent allows developers to run complex, multi-iterion agent pipelines at a fraction of the standard API cost.
Although Hy4 boasts 770 billion total parameters, Mixture-of-Experts (MoE) routing enables only 49 billion parameters per token. This selective enablement strategy keeps inference latency manageable. And by balancing GPU memory bandwidth requirements, tokens can be generated quickly even when processing long and contextually complex alert messages.
Source: Tencent Hy
Comments
Post a Comment