AI × Drug Discovery Research profile

Tianxiang Wu

吴天翔
吴天翔 · Tianxiang Wu · wutianxiang

Building reliable scientific data systems and representation learning methods for computational drug discovery.

Formal portrait of Tianxiang Wu
Academic profile Tianxiang Wu / 吴天翔

Computer Science undergraduate · AI for Drug Discovery

National Scholarship Bioinformatics · CCF A Rank 2 / 96
CurrentTsinghua AIRResearch Intern
EducationXidian UniversityComputer Science
FocusAI4ScienceData · Molecules · Proteins
Scroll to research
00

Academic snapshot

An academic profile across AI for Science.

Computer Science undergraduate at Xidian University with research training across AIDD agents, ADMET data systems, molecular representation learning and peptide modeling.

GPA3.9 / 4.0Core average 92+
Rank2 / 96Track ranking
EnglishCET-6 555Literature reading
ScholarshipNational2024-2025
Honors

Scholarships & competitions

  • National Scholarship for undergraduates, 2024-2025.
  • Lenovo Scholarship, 2025-2026, the only recipient in the grade.
  • Xinghuo Cup LLM Agent track university first prize.
  • Xinghuo Cup academic A-track school preliminary special prize and first prize.
Outputs

Papers & manuscripts

  • Bioinformatics work on peptide engineering language models, published as second author / undergraduate lead author.
  • 3D-MPG molecular pretraining manuscript submitted to AAAI 2027, a CCF A conference, as second author / undergraduate lead author.
  • Peptide property prediction benchmark manuscript submitted to Briefings in Bioinformatics.
Identity & service

Leadership, public work & arts

  • Probationary CPC member; Outstanding Communist Youth League Member Model, one of ten across the university.
  • Community second-level grid worker, with 714 hours of volunteer service across community, education and public-service activities.
  • Cambridge summer visiting program project lead, Xidian Symphony Orchestra violinist, and student committee service in class.
AI for Drug Discovery Scientific Information Extraction ADMET Data Infrastructure Molecular Representation Learning Protein Modeling
01

Selected research

From messy evidence
to useful models.

Current work centers on scientific data extraction, context-aware molecular modeling, and computational drug discovery.

02 / Condition-aware predictionOngoing

Unified two-tower
ADMET modeling

Designing a unified two-tower ADMET prediction framework that treats assay and experimental context as first-class model inputs rather than metadata left outside the model.

ADMETAssay ContextMultitask
View model work
03 / Peptide engineering language modelResearch

Instruction-tuned
peptide LLM

Built and trained a peptide instruction model with LoRA, covering function description, sequence design, property prediction and physicochemical optimization.

LoRAPeptideMultitask
Open paper
04 / 3D molecular pretrainingResearch

3D-MPG geometry
pretraining

Contributed to 3D-MPG, a molecular representation learning manuscript on relational geometry pretraining through intramolecular region-pair arrangements, submitted to AAAI 2027.

3D GeometryEGNNAAAI 2027
View submission record
05 / Task-adaptive molecular platformResearch

Molecular pretraining
platform

Built a task-adaptive molecular pretraining platform recognized with Xinghuo Cup university first prize and college-level special prize.

Molecular PretrainingData QualityAwarded
View award record
06 / Drug-discovery agentsExploration

PharmAgent
research workflows

Exploring PharmAgent workflows that connect autonomous research planning, molecular evidence and downstream drug-discovery tasks.

Agent WorkflowLLM AgentsDrug Discovery
View agent systems
07 / Protein-ligand foundation modelOngoing

Dual-tower
foundation model

Improving the sequence tower of a DrugCLIP-style protein-ligand foundation model with ESM-C representations and hyperbolic latent structure for OOD generalization across targets, scaffolds and low-data settings.

DrugCLIPESM-COOD
View ADMET work
02

Publications & outputs

Published and submitted
research outputs.

Selected published and submitted work is presented with current status, authorship role and supporting records.

View published paper
Published · Bioinformatics · CCF A

Bridging Linguistic Reasoning and Biophysical Reality Toward Peptide Engineering via Instruction-Tuned Language Modelling

Second author / undergraduate lead author. The work connects instruction-tuned language modeling with peptide engineering tasks and was published in Bioinformatics.

Submitted · AAAI 2027 · CCF A

3D-MPG: Relational Geometry Pretraining through Intramolecular Region-Pair Arrangements

Second author / undergraduate lead author. The manuscript focuses on 3D molecular geometry pretraining for molecular representation learning.

View submission record
AAAI submission evidence
Submitted · Briefings in Bioinformatics · JCR Q1

A Systematic Benchmark for Peptide Property Prediction

Third author. The manuscript builds a systematic benchmark for peptide property prediction.

View manuscript record
Peptide benchmark manuscript evidence
03

Life moments

Research, service
and life in motion

Research training, academic sharing, international exchange and orchestra performance

Arts record Xidian Symphony Orchestra member

Member of Xidian Symphony Orchestra, a Shaanxi high-level arts troupe and first-chair troupe of Shaanxi collegiate arts troupe; participated in major performances including the ICAI 2024 opening ceremony.

Shaanxi high-level arts troupe First-chair collegiate arts troupe ICAI 2024 opening ceremony
04

About

I care about one question:

How do we make AI systems scientifically useful, not just benchmark-good?

I am a Computer Science undergraduate at Xidian University and a research intern at the Institute for AI Industry Research (AIR), Tsinghua University.

My current interests include AI for Drug Discovery, patent-scale scientific information extraction, high-quality ADMET data construction, and representation learning for molecules and proteins.

I am especially interested in systems that connect robust data infrastructure with predictive models that remain useful outside curated benchmarks.

01Evidence first

Predictions should remain traceable to scientific evidence.

02Context matters

Model assay and experimental context, not only molecules.

03Quality is research

Data quality is part of the scientific method, not cleanup.

05

Experience

Academic trajectory.

Computer science → AI4Science → drug discovery.

CURRENT
Tsinghua University · Institute for AI Industry Research

Research Intern

AI for Drug Discovery, scientific data extraction, and autonomous research systems.

01
UNDERGRAD
Xidian University

B.Sc. · Computer Science & Technology

Computer science training with a research focus on AI for scientific discovery.

02
Research · Collaboration · Academic exchange

Find me
online.