General

Associate Professor

zhangheng17@iscas.ac.cn

Institute of Software, 

Chinese Academy of Sciences (ISCAS)

Research Areas

My current research centers on embodied intelligence, operating systems, and high‑performance computing systems. In this vein, I have led a cross‑functional team to develop Alpha Platform—an integrated development environment that unifies simulation, distributed training, and real‑world deployment for embodied AI agents. Unlike isolated toolchains, Alpha Platform provides a seamless pipeline from rapid algorithm prototyping to large‑scale reinforcement learning on GPU clusters, and further supports automated sim‑to‑real transfer with domain randomization and latency‑aware adaptation. The platform has been successfully applied to quadruped locomotion, manipulation tasks, and multi‑agent coordination scenarios, reducing development cycles in internal evaluations.


Alongside this platform effort, I have directed the design and implementation of multiple high‑performance system solutions that span both conventional x86 architectures and emerging accelerator‑based platforms. My broader research agenda pursues three interconnected threads: (1) building scalable, high‑throughput systems for large‑scale data analytics; (2) exploring software‑hardware co‑design opportunities on novel accelerators (GPUs, DPUs, FPGAs); and (3) tackling modern computing challenges in datacenter‑scale AI, intelligent edge, and next‑generation HPC environments. My contributions in these areas range from low‑level system optimizations—such as GPU‑aware framework acceleration, tiered memory management, and computation‑aware task scheduling—to the design of adaptive parallel paradigms that dynamically adjust to the algorithmic structure and runtime workload characteristics of diverse applications. I also possess substantial hands‑on experience in constructing full software stacks, high‑performance numerical libraries (e.g., custom BLAS/LAPACK variants), and turnkey big‑data platforms that significantly lower the barrier for end‑user deployment.


At present, my focused research efforts are directed toward: high‑performance frameworks for heterogeneous architectures (GPU, DPU, FPGA), distributed continual learning frameworks that exhibit near‑linear scalability up to thousands of nodes, robotic systems with real‑time and safety‑critical guarantees, and agent memory techniques that effectively bridge the performance gap between DRAM and non‑volatile storage. Moreover, I am actively investigating how these system‑level innovations can be synergized to accelerate embodied AI, particularly in the areas of simulation‑to‑real generalization, low‑latency closed‑loop control, and resource‑efficient decision making under physical constraints.


Previously, I received my Ph.D. degree from the Institute of Software, Chinese Academy of Sciences, in January 2018. My doctoral research, which focused on adaptive runtime systems for irregular applications, laid the foundation for my enduring interest in system performance, programmability, and hardware‑aware design. Since then, I have consistently strived to push the frontiers of system efficiency and ease‑of‑use, delivering practical, deployable solutions that keep pace with the rapidly evolving demands of AI workloads and next‑generation hardware architectures.

Education

  • Sept. 2012 – Mar. 2018  Institute of Software, Chinese Academy of Sciences, Beijing, China 

            Ph.D student in Computer Software and Theory

  • Sept. 2008 – Jun. 2012  Northeastern University

            B.S. in Computer Software Engineering

Experience


Work Experience

May 2018 – Now, Institute of Software, Chinese Academy of Sciences (ISCAS), Beijing, China 

Associate Research Professor, focus on high perofmance computing / systems / machine learning


Publications


Papers

[1]Kaifan Jia, Yongchun Jiang, Zhihao Ling, Minghui Zhang, Xuran Wang, Ran Bao, Haonan Zou, Heng Zhang*. BTC-TC: Exact GPU Triangle Counting with Hybrid Bit Tensor Cores and CUDA Cores. The International Conference for High Performance Computing, Networking, Storage, and Analysis (SC), 2026. 

[2]Yongchun Jiang, Yongchao Liu, Kaifan Jia, Heng Zhang*. ATLAS: Enabling an Asynchronous State-Aware Pipeline for Efficient Training of Temporal Graph Neural Networks. The 55th International Conference on Parallel Processing (ICPP 2026).

[3]Zhirui Chen, Heng Zhang*, Kaifan Jia. AFH-SpMM: Auto-Fit Heterogeneous Block Sparse-Dense Matrix Multiplication on Tensor Core GPUs. The 55th International Conference on Parallel Processing (ICPP 2026).

[4]Xiangfei Fang, Ran Bao, Heng Zhang*. CaN: A Core-aware Neural Framework for Attributed Hypergraph Generation. 32nd SIGKDD Conference on Knowledge Discovery and Data Mining (KDD), 2026.

[5]YongChun Jiang, Heng Zhang*, Jian Gao, Xin Zheng. History Doesn’t Repeat, But Its Patterns Echo: A Parallel Pairwise Negative-Sampling Framework for Temporal Link Prediction. 35th International Joint Conference on Artificial Intelligence (IJCAI), 2026.

[6]Haonan Zou, Heng Zhang*, Yongchun Jiang, Kaifan Jia, Yansong Dong, Shang Wang and Zhihao Ling. DATEE: An Adaptively Thresholded Early-Exit Framework for Large Language Model Inference Based on Confidence Trend Sampling. The 19th International Conference on Knowledge Science, Engineering and Management (KSEM 2026).

[7]Shang Wang, Yansong Dong, Kaifan Jia, Haonan Zou and Heng Zhang*. xHyperG: A Hypergraph Analytical Framework on GPUs with Scalability, in 2025 IEEE 31th International Conference on Parallel and Distributed Systems (ICPADS), 2025, pp. 1-9.

[8]Peiyuan Dai, Rui Liu, Heng Zhang*. NanoCSV: Enabling Efficient Parallel CSV Extraction with Hierarchical Finite-State Transducer[C]//International Conference on Database Systems for Advanced Applications (DASFAA). Singapore: Springer Nature Singapore, 2025.

[9]Xiangfei Fang, Chengying Huan, Heng Zhang, Yongchao Liu, Shaonan Ma, Yanjun Wu, Chen Zhao.. OTM: Efficient K-Order-Based Core Maintenance in Large-Scale Dynamic Hypergraphs[J]. IEEE Transactions on Knowledge Discovery from Data (TKDD), 2025.

[10]Xiangfei Fang, Chengying Huan, Boying Wang, Shaonan Ma, Heng Zhang, Chen Zhao. HyperSF: A Hypergraph Representation Learning Method Based on Structural Fusion. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Hyderabad, India, 2025

[11]Xiangfei Fang, Boying Wang, Chengying Huan, Shaonan Ma, Heng Zhang, Chen Zhao. HyperKAN: Hypergraph Representation Learning with Kolmogorov-Arnold Networks. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Hyderabad, India, 2025

[12]Chengying Huan, Heng Zhang*, Likang Chen, Yongchao Liu, Xuran Wang, Yongchun Jiang, Shaonan Ma, Yanjun Wu. TeMatch: A Fast Temporal Subgraph Matching Framework with Temporal-Aware Subgraph Matching Algorithms. IEEE International Conference on Data Engineering (ICDE). 2025

[13]Qinglin Pan, Ji Qi, Jiatai He, Heng Zhang, Jiageng Yu, Yanjun Wu, Beaver: A High-Performance and Crash-Consistent File System Cache via PM-DRAM Collaborative Memory Tiering. International Conference on Measurement and Modeling of Computer Systems (SIGMETRICS), 2025.

[14]Huan, Chengying and Liu, Yongchao and Zhang, Heng* and Liu, Hang and Chen, Shiyang and Song, Shuaiwen Leon and Wu, Yanjun. TeGraph+: Scalable Temporal Graph Processing Enabling Flexible Edge Modifications [J]. IEEE Transactions on Parallel and Distributed Systems (TPDS), 2024. doi: 10.1109/TPDS.2024.3393914. 

[15]Chengying Huan, Yongchao Liu, Heng Zhang*, Shuaiwen Leon Song, Shiyang Chen, Xiangfei Fang, Yue Jin, Baptiste Lepers, Yanjun Wu, Hang Liu. TEA+: A Novel Temporal Graph Random Walk Engine With Hybrid Storage Architecture [J]. ACM Transactions on Architecture and Code Optimization (TACO), 2024.

[16]Yue Jin, Chengying Huan, Heng Zhang, Yongchao Liu, Shuaiwen Leon Song, Rui Zhao, Yao Zhang, Changhua He, Wenguang Chen. G-Sparse: compiler-driven acceleration for generalized sparse computation for graph neural networks on modern GPUs. 32nd International Conference on Parallel Architectures and Compilation Techniques (PACT 2023), 2023

[17]Heng Zhang, Lingda Li, Hang Liu, Dongling Zhuang, Rui Liu, Chengyin Huan, Charles He, Yongchao Liu, Shuang Song, Dingwen Tao, Yanjun Wu, Shuaiwen Song. Bring Orders into Uncertainty: Enabling Efficient Uncertain Graph Processing via Novel Path Sampling on Multi-Accelerator System [C]. ACM International Conference on Supercomputing (ICS), 2022. ICS’22.

[18]Chengying Huan, Shuaiwen Leon Song, Yongchao Liu, Heng Zhang, Hang Liu, Charles He, Kang Chen, Jinlei Jiang, Yongwei Wu. T-GCN: A Sampling Based Streaming Graph Neural Network System With Hybrid Architecture [C]. 31st International Conference on Parallel Architectures and Compilation Techniques (PACT 2022).

[19]Heng Zhang, Lingda Li, Dongling Zhuang, Rui Liu, Shuang Song, Dingwen Tao, Yanjun Wu, Shuaiwen Song. An Efficient Uncertain Graph Processing Framework for Heterogeneous Architectures[C]. ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (PPoPP), 2021. PPoPP’21.

[20]Pengpeng Hou, Heng Zhang, Yanjun Wu, Jiageng Yu, Yang He, Yuxia Miao. FindCmd: A Personalized Command Retrieval Tool[J]. IET Software, 2021.

[21]Hongjun Zhang, Heng Zhang, Libo Zhang, Yanjun Wu. FastUDP: A highly scalable user-level UDP framework in multi-core systems for fast packet I/O[J]. The Journal of Supercomputing, 2020, 11(3).

[22]Heng Zhang, Libo Zhang, Da Cheng, Yanjun Wu, Chen Zhao. EpCom: A parallel community detection approach for epidemic diffusion over social networks[C]. 2017 IEEE International Conference on Bioinformatics and Biomedicine. IEEE, 2017: 1607-1614. BIBM 2017.

[23]Heng Zhang, Haibo Hou, Libo Zhang, Hongjun Zhang, Yanjun Wu. Accelerating Core Decomposition in Large Temporal Networks Using GPUs[C]. International Conference on Neural Information Processing. Springer, Cham, 2017: 893-903. ICONIP’17.

[24]Heng Zhang, Chunliang Hao, Yanjun Wu, Mingshu Li. Towards a scalable and energy-efficient resource manager for coupling cluster computing with distributed embedded computing[J]. Cluster Computing, 2017, 20(4): 3707-3720. Cluster Computing, 2017. 

[25]Heng Zhang, Chunliang Hao, Yanjun Wu, Mingshu Li. Macaca: a scalable and energy-efficient platform for coupling cloud computing with distributed embedded computing[C]. IEEE Parallel and Distributed Processing Symposium Workshops. IEEE, 2016: 1785-1788. IPDRM’16, IPDPS 2016.

[26]Chunliang Hao, Jie Shen, Celia Chen, Heng Zhang, Yanjun Wu, Mingshu Li. PCSsampler: Sample-based, Private-state Cluster Scheduling[C]. Proceedings of the 17th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing. IEEE Press, 2017: 599-608. CCGrid 2017.

[27]Chunliang Hao, Jie Shen, Heng Zhang, Xiao Zhang, Yanjun Wu and Mingshu Li. Tiresias: low-overhead sample based scheduling with task hopping[C]. IEEE International Conference on Cluster Computing. IEEE, 2016: 251-254. IEEE Cluster 2016. 

[28]Chunliang Hao, Jie Shen, Heng Zhang, Xiao Zhang, Yanjun Wu and Mingshu Li. Sparkle: adaptive sample based scheduling for cluster computing[C]. Proceedings of the 5th International Workshop on Cloud Data and Platforms. ACM, 2015: 5. CloudDP 2015, Eurosys 2015

[29]Kai Wang, Heng Zhang, Yanjun Wu, Chen Zhao, Mingshu Li. Frog: A Distributed Graph Processing Engine from Sequential Subgraph Blocks. Work-in-Progress, ACM SOSP 2013.

Research Interests

High Performance Computing

Operating System

Distributed and Parallel System

Machine Learning

Collaboration

Jan. 2021-Jan. 2022, University of Sydney, Austrilia

Research Visting Scholar, FSA Lab (Professor Shuaiwen Leon Song), University of Sydney, Austrilia

• Work on high performance system building and big-data analytics.


Nov. 2014 – Jun. 2015, Intel Lab China, Beijing

Research Internship, Data Infrastructure Laboratory (DIL) in Intel Lab China, Intel Inc.

• Work on distributed communication framework optimization, e.g., Petuum system, Memcached, etc.


Nov. 2011 – Apr. 2012, Baidu Inc., Beijing

Software Engineering Internship, Baidu Map Team 

• Responsible for launching webmap 2.0 online with team and webmap API v1.3 maintenance.

Students

已指导学生

代培元  硕士研究生  085405-软件工程  

董岩松  硕士研究生  085405-软件工程  

王上  硕士研究生  085405-软件工程  

邹浩南  硕士研究生  085410-人工智能  

现指导学生

姜永春  硕士研究生  085405-软件工程  

包冉  硕士研究生  083500-软件工程  

凌志豪  硕士研究生  083500-软件工程  

王旭然  硕士研究生  083500-软件工程  

夏宇航  硕士研究生  085405-软件工程  

朱小懋  硕士研究生  083500-软件工程  

张明辉  硕士研究生  085410-人工智能