Senior Applied Scientist, Document Understanding

Thomson Reuters - Eagan, MN

Hiring: Senior Applied Scientist, Document Understanding Company: Thomson Reuters Location: Eagan, MN Job Posted Time: 2026-09-16 11:20:04 Target Skills & Keywords : AI, AWS, CAD, Embedded Systems, Hugging Face, LLM, NLP, PyTorch, Python, RAG, SageMaker, Transformers About the job Experience: •5+ years of post-degree industry experience shipping document understanding, information extraction, or knowledge graph systems into production — not research-only experience Required Skills: •This is an applied science position focused on designing, building, and deploying production-grade document understanding systems that power Westlaw, PracticalLaw, and CoCounsel. •Work across semantic chunking, document enrichment, and knowledge graph construction for complex legal, tax, and accounting content — delivering foundational intelligence that multiple product teams depend on at scale. •Design and deploy semantic chunking models for lengthy, non-uniformly structured legal documents with adjustable granularity across use cases •Build document enrichment systems that classify documents according to legal and customer-defined taxonomies and extract rich metadata •Develop LLM-based knowledge graph construction pipelines that extract and link citations, entities, and legal concepts across diverse legal content •Build scalable synthetic data generation systems for model training, multi-hop query simulation, and hallucination-free answer generation •Design evaluation frameworks — component-level and end-to-end — using expert annotation and synthetic data •Drive independent technical decisions on chunking strategy, classification approach, knowledge extraction methods, and multi-document reasoning architecture Qualifications: •PhD or Master's in Computer Science, AI, NLP, or a related field •Publications at ACL, EMNLP, ICLR, NeurIPS, SIGIR, KDD, or equivalent •Production Python and experience with PyTorch, Hugging Face Transformers, and DeepSpeed •Hands-on Production Depth Required In Document layout analysis and semantic chunking beyond fixed-size or paragraph-based methods •Hierarchical, multi-label document classification with domain-specific and customer-defined schemas •Entity recognition and linking, relation extraction, citation parsing, and knowledge graph construction from unstructured text •LLM-based information extraction, few-shot and multi-task learning, and post-training •Knowledge distillation, model compression, and SLM deployment under latency constraints •Synthetic data generation for NLP: query-answer generation with verification and scalable data augmentation •Annotation workflow design and evaluation framework development for document understanding tasks Compensation: •Flexible work environment (work from home / hybrid options) •Competitive benefits and rewards package Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!