Futong CloudOne × Futong TokenWise: Building an Innovative Provincial Medical AI Computing Power Operations Hub
With the continued advancement of the “AI + Healthcare” initiative, AI technologies such as large language models (LLMs) and intelligent agents are rapidly being introduced into application scenarios including clinical assistance, medical research, and health management. This has placed higher requirements on underlying computing resources, model service capabilities, and operational management systems.
During the development of regional medical AI capabilities, the Health Commission of a province initiated the construction of an AI computing power operations hub, aiming to integrate scattered computing resources and AI capabilities across the region and provide unified, efficient, and secure AI infrastructure services for medical institutions, research organizations, and other users.
To support this objective, Futong Technology (00465.HK), leveraging the capabilities of its two core products — Futong CloudOne Enterprise Digital Full-Asset Operations Platform and Futong TokenWise Token Operations Platform — undertook the construction of this provincial medical AI computing power operations hub project. The project enabled unified management of multi-source computing resources, standardized access to model services, refined Token usage measurement, and visualized platform operations management, exploring a new service-oriented operations model for AI infrastructure in the healthcare sector.
This practice was successfully selected as one of the 2026 Trusted Token Cloud Service Innovation Practices.
Project Practice: Addressing Three Key Business Challenges in Scaling Medical AI Operations
Before the project started, the province had already deployed multiple AI computing resources to provide foundational computing support for medical institutions and research organizations. However, as large language models and intelligent agent applications gradually entered business scenarios, customers found that the key constraint for large-scale AI adoption was no longer resource construction capability, but rather operational capability.
These challenges were mainly concentrated in three core business scenarios: computing resource supply, AI service operations, and healthcare data security.
Scenario 1: Regional Computing Resources Are “Unevenly Utilized” — How Can Computing Resources Respond Dynamically to Demand?
Typical Business Challenges
Although the overall amount of computing resources was not insufficient, resources were often “unavailable when needed.” Multiple types of computing resources, including self-operated cloud, partner-operated cloud, and third-party public cloud resources, were independently built and operated under different systems. As a result, resource status lacked transparency, and operations teams were unable to obtain real-time visibility into computing resource distribution and utilization across the entire region.
This situation directly affected medical institutions and research organizations. When clinical departments or research teams submitted computing resource requests, operators needed to check availability across multiple resource pools individually, resulting in lengthy processes and uncertain outcomes. Some research projects were delayed due to long resource delivery cycles, while the deployment efficiency of AI applications in some medical institutions was constrained. Meanwhile, other resource pools remained idle for extended periods, with overall GPU utilization remaining below 40%.
The fundamental issue was not a shortage of computing resources, but the lack of unified operations and scheduling mechanisms. This resulted in a structural mismatch between supply and demand, making it difficult to establish sustainable, stable, and efficient computing service capabilities for medical institutions across the province.
Futong Solution
Futong CloudOne helped the customer establish a unified computing resource operations system by integrating and centrally managing distributed computing resources across multi-source heterogeneous environments, transforming them into a unified computing resource pool that could be centrally allocated.
The platform encapsulates underlying GPU/NPU resources into standardized computing services. When business departments submit computing requirements, operations teams can complete resource matching and allocation through a one-stop process on the platform without needing to know the underlying chip type or cloud platform source.
When medical institutions or research teams submit new computing requests, operations teams can view available resource status in real time through the platform and complete computing resource allocation and delivery based on actual demand.
The time required to onboard new computing resources into the platform has been reduced from an average of two weeks to three days. Previously fragmented computing resources can now be centrally scheduled and dynamically allocated based on demand, becoming a truly shared regional infrastructure capability.
Scenario 2: AI Usage Continues to Grow — How Can Token Consumption Be Accurately Measured and Clearly Allocated?
Typical Business Challenges
As the regional AI computing power hub entered operation, more hospitals, research institutions, and business departments began accessing and using model services provided by the platform. AI service requests continued to grow rapidly, with the daily peak number of AI calls exceeding 2 million.
Behind this growth, a new operational challenge emerged: every AI request generates Token consumption, and Token consumption directly corresponds to cost. However, without a refined measurement system, platform operators faced a series of unavoidable questions — which tenants, departments, and projects were using AI services? How many Tokens were consumed by each? And which entities should bear the corresponding costs?
The impact of this issue was multi-dimensional. During internal settlement, inaccurate usage data could lead to disputes; budget management lacked reliable consumption data for precise planning; and operators were unable to provide trustworthy billing statements to users. Without unified measurement, cost accounting, and operational management capabilities, large-scale AI service operations lacked a solid foundation.
Before the project implementation, the actual Token undercounting rate on the regional platform reached approximately 5%, meaning that a significant portion of AI services had already been consumed but was not accurately recorded in cost accounts.
Futong Solution
Futong TokenWise established a request-level refined Token operations system. For every AI request, the platform records key information including the requester, request time, Token consumption amount, and model used. This enables precise attribution of Token consumption to specific users and business scenarios, supporting comprehensive operational capabilities including consumption statistics, cost allocation, and account reconciliation.
To ensure that every usage record could be accurately captured under high-concurrency scenarios without affecting business response performance, the project adopted an “edge-side pre-aggregation + asynchronous reporting” approach. Edge nodes perform real-time pre-aggregation of second-level data, while central nodes asynchronously persist the data, achieving a balance between real-time performance (less than 3 seconds latency) and accuracy (99.99%).
Every AI request becomes traceable and auditable, providing a reliable measurement foundation for commercial operations.
Scenario 3: Healthcare Data Cannot Leave Its Domain — How Can AI Services Truly Reach Business Scenarios?
Typical Business Challenges
The healthcare industry has strict requirements for patient data security due to regulatory requirements and industry compliance standards. For hospitals and information management authorities, original medical data such as patient imaging, medical records, and test reports must never be uploaded to public clouds or external platforms under any circumstances. This represents a non-negotiable security boundary.
However, AI models within regional AI computing power hubs are typically deployed at central nodes, while data remains distributed within individual hospitals. This creates a fundamental challenge: with models located centrally and data retained inside hospitals, how can AI services securely reach real healthcare business scenarios?
If hospitals were required to upload data due to AI service requirements, compliance requirements could not be met. However, if the use of centralized AI model capabilities was abandoned due to security concerns, the value of building the regional AI hub would be significantly reduced.
Futong Solution
The project adopted a “data remains in place, model moves” technical approach.
The platform does not collect or store any original medical data. Instead, it focuses only on model scheduling and operational management. When a hospital initiates an AI inference request, the platform delivers the required model to the hospital’s local computing nodes through dedicated network connections. The inference process is completed entirely within the hospital, and inference results are returned to the platform after desensitization. The platform only retains request logs and billing information.
Meanwhile, the platform established a three-layer security framework covering network, application, and data security, enabling capabilities including tenant isolation, RBAC-based access control, and full-process operational auditing. These capabilities comprehensively meet the Level 3 requirements of China’s cybersecurity multi-level protection compliance for healthcare scenarios.
With security and compliance ensured, the AI service capabilities of the hub can truly reach core healthcare business scenarios.
CloudOne Builds the Foundation, TokenWise Enables Operations: Jointly Building an AI Infrastructure Operations System
Based on Futong CloudOne’s digital operations hub capabilities and Futong TokenWise’s Token operations capabilities, Futong Technology built an AI computing power operations architecture for this project covering resource supply, model services, Token operations, security governance, and operational visualization.
Resource Operations Layer (Futong CloudOne):
Responsible for unified access, centralized management, and unified scheduling of multi-source heterogeneous computing resources, enabling pooled operations of computing resources. At the same time, it provides users with a unified service portal, supporting end-to-end services including computing resource requests, model invocation, and usage queries.Model Service and Token Operations Layer (Futong TokenWise):
Enables standardized access to multiple model services, intelligent routing and scheduling, as well as request-level Token measurement, cost accounting, and operational analysis, making AI services measurable, attributable, and accountable.Security Compliance and Operational Visualization Layer:
Security and compliance capabilities run throughout resource, model, and operations layers. Through visualization modules such as computing center distribution maps, resource monitoring dashboards, and operational overview dashboards, the platform enables transparent management of AI infrastructure operating status.
Through coordinated operation of these layers, the project establishes a complete system supporting sustainable operations of regional medical AI computing infrastructure.
Project Outcomes: Enabling AI Computing Power to Truly Become a Service-Oriented Operation
Currently, the platform has completed unified management of multiple computing centers, forming a hundred-PFLOPS-level intelligent computing resource pool. It provides AI computing services for more than 20 hospitals and over 10 research institutions, with platform operational data continuously displayed through visualization dashboards.
At the Resource Operations Level, the Customer Established a Regional Unified Computing Resource Operations System for the First Time, Transforming Resource Management from Fragmented Management to Centralized Operations:
Overall GPU resource utilization increased from 40% to 70%.
The multi-console operations and maintenance model was upgraded into a unified operations portal, reducing the average fault identification time from hours to minutes.
The average onboarding cycle for new models and new computing resources was shortened from two weeks to three days.
At the AI Operations Level, the Customer Further Established a Continuous AI Operations System Covering Model Access, Token Measurement, Cost Accounting, and Ecosystem Enablement:
Based on request-level Token measurement capabilities, the Token consumption generated by every AI request can be accurately attributed to tenants, departments, and projects, enabling refined AI cost accounting. The resource undercounting rate was reduced from approximately 5% to below 0.1%.
The unified API service system lowered the entry barriers for third-party developers and partners. Currently, more than 10 healthcare AI companies have conducted application development based on the platform, with multiple healthcare models and AI applications deployed, initially forming a regional AI computing service ecosystem.
As AI infrastructure development gradually enters a stage of large-scale operations, the focus of enterprise competition is shifting from “capability building” to “operational capability.”
This project not only demonstrates Futong Technology’s capabilities in building and operating complex AI infrastructure environments in the healthcare sector, but also showcases the practical value of jointly building an AI infrastructure operations system through Futong CloudOne and Futong TokenWise.
This capability system can also be further extended to more industry scenarios, including government services, scientific research, and manufacturing. It will continue to support enterprises and industry customers in building unified, efficient, secure, and operable AI infrastructure, providing a solid foundation for the large-scale adoption of artificial intelligence.