Principal Applied Scientist,
AWS Agentic AI
Haofu Liao 廖昊夫
I am a tech lead and Principal Applied Scientist at AWS Agentic AI, and previously led document-understanding research at AWS AI Labs. I received my Ph.D. in Computer Science from the University of Rochester, where I worked with Jiebo Luo, my M.S. in Electrical and Computer Engineering from Northeastern University in Boston, and my B.Eng. from the Beijing University of Posts and Telecommunications.
My work centers on agentic AI: building autonomous agents and the foundation models behind them that can interpret multimodal inputs, reason over complex context, and act to complete real-world tasks. My research spans agentic systems, large action models, and the orchestration that turns models into reliable agents.
At AWS, I build the agentic systems and train the large action models behind Amazon Quick, where I lead the science for its workflow automation (Quick Automate and Quick Flows). Before that, I led the science behind AWS products including Bedrock Data Automation and Textract, and developed document-understanding methods such as DocTr and DocKD. My earlier research in medical image computing introduced methods like DuDoNet and ADN for CT artifact reduction. Altogether, my work has been cited over 2,600 times with an h-index of 26.
I typically host one summer intern each year, so if you are interested in an internship with AWS Agentic AI, feel free to reach out.
News
- 06/2026AWS DevOps Agent added release management capabilities in preview, including release readiness review and autonomous release testing before production.
- 04/2026I was promoted to Principal Applied Scientist at AWS Agentic AI.
- 02/2026Our paper on vision-language diffusion models for GUI grounding was accepted to CVPR 2026.
- 10/2025Amazon Quick launched, an agentic AI platform for workplace automation. I lead the science for its agentic workflow automation.
- 05/2025Our paper on efficient web automation was accepted to Findings of ACL 2025.
- 03/2025Amazon Bedrock Data Automation became generally available. I contributed to the service as a tech lead for model training through public preview.
- 09/2024Our DocKD paper was accepted to EMNLP 2024.
- 05/2024I moved to AWS Agentic AI to lead the science for agentic workflow automation.
Selected Research
See all research →
Amazon Quick Workflow Automation
I lead the science for agentic workflow automation in Amazon Quick, AWS's agentic AI platform for workplace automation. I started with GUI automation, building a vision-language large action model from scratch in 2024 when no comparable solution existed in the industry, and directed teams across data, mid-training, and post-training to take it from research to production. I have since expanded to lead the science across the full stack of workflow automation, where agents complete tasks over user interfaces, APIs, tools, and code, spanning Quick Automate for complex multi-step processes and Quick Flows for no-code automation of routine tasks. This work is reflected in our research on GUI grounding (CVPR 2026) and efficient web automation (Findings of ACL 2025).
Quick AutomateGUI Grounding (CVPR 2026)Web Automation (ACL 2025)

AWS DevOps Agent Release Management
I am the tech lead for autonomous release testing in AWS DevOps Agent's release management capabilities. The agent generates and runs change-specific tests for web and API applications in production-like environments, helping teams assess functional correctness, behavioral regressions, and integration risks before changes reach production.

Amazon Bedrock Data Automation
I was a tech lead for Amazon Bedrock Data Automation, the AWS service that turns unstructured multimodal content (documents, images, video, and audio) into structured insights for generative-AI applications, where I led model training for several core components and helped bring the service to public preview. On the problem of making compact document-understanding models generalize to unseen formats, I co-led DocKD (co-first author, EMNLP 2024), a knowledge-distillation method that feeds large language models structured document elements (key-value pairs, layouts, and descriptions) to generate higher-quality synthetic training data. Models trained only on DocKD data match human-annotated data in-domain and surpass it out-of-domain.

Amazon Textract Analyze Expense
I was a tech lead for Amazon Textract Analyze Expense, the capability that extracts fields and line items from invoices and receipts without templates. On this structured-extraction problem I led DocTr (first author, ICCV 2023), a document transformer that recasts information extraction as anchor-based detection: each entity is represented by an anchor word and a bounding box, with relations captured through anchor-word associations, replacing the fragile token tagging and brittle graph decoding of prior methods. DocTr achieved state-of-the-art results across three extraction benchmarks.
Selected Publications
See all publications →
Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026.

Turbocharging Web Automation: The Impact of Compressed History States
Findings of the Association for Computational Linguistics (ACL), 2025.

DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding Models
Conference on Empirical Methods in Natural Language Processing (EMNLP), 2024.

DocTr: Document Transformer for Structured Information Extraction in Documents
IEEE/CVF International Conference on Computer Vision (ICCV), 2023.

Deep Network Design for Medical Image Computing: Principles and Applications
Academic Press, 2022. Book