Skip to content

AI Data Engineer (Materials Data Extraction)

  • On-site
    • Singapore, Central Singapore, Singapore
  • SG - Materials

Job description

Data Engineer / Senior Data Engineer

Patsnap is looking for a Data Engineer or Senior Data Engineer to help build and evolve the data foundation behind our global innovation intelligence and AI-powered products. You’ll work with large-scale patent and related datasets, designing robust systems that transform complex raw information into accurate, reliable, and accessible data that our products and customers can depend on.

This is a hands-on engineering role with ownership across the full data lifecycle—from ingestion and parsing through to storage, quality, performance, and availability. You’ll solve challenging data problems at scale, continuously improve the efficiency and accuracy of our processing frameworks, and work closely with data architects to evolve our platform for greater scalability and resilience. Your work will directly enable Patsnap to deliver trusted, high-quality data and rapidly bring new data capabilities to customers around the world.

Want to see the platform you'd be representing?
Check out this short overview:

This is an in-office position based in our Singapore office.

 

Who are we?

Patsnap is a global, pre-IPO company that transforms the way organizations harness their Intellectual Property and Research & Development productivity. Our platform revolutionizes how IP and R&D teams collaborate across the entire innovation lifecycle, using domain-specific AI to accelerate the creation of market-ready products. With over 12,000 customers worldwide, including some of the biggest names in innovation, Patsnap is at the forefront of technological advancement. Our $300M Series E funding round brings our valuation to a $1 billion unicorn status, and we still have a remarkable amount of growth ahead. 

We have a vibrant and diverse team with offices in Singapore, Toronto, London, Shanghai and remote teams based in US. Our hyper-growth trajectory is powered by our people, and we are extremely proud of our company-wide vision, work ethic, and entrepreneurial spirit. We are committed to fostering an inclusive environment where talent thrives and ideas bloom.   

What You'll Be Doing:

  • Process and extract structured information from materials-related patents, scientific literature, and technical documents.

  • Design and test prompts, input formats, and extraction workflows across different models and datasets.

  • Build LLM workflows covering input preparation, model calls, prompt versioning, validation, evaluation, retries, and human review.

  • Analyse model failures and improve results through changes to prompts, rules, data, or workflow design.

  • Help build gold-standard datasets, annotation guidelines, quality checks, and automated evaluations.

  • Use Python and SQL for data cleaning, transformation, batch processing, and result analysis.

  • Improve the accuracy, reliability, processing speed, and cost efficiency of data extraction workflows.

  • Turn successful experiments into reusable scripts, tools, skills, or automated workflows.

Job requirements

  • Bachelor’s degree in Computer Science, Software Engineering, Data Science, Artificial Intelligence, Materials Science, or a related discipline.

  • Good programming fundamentals, with the ability to use Python or Java for basic development and data processing.

  • Basic knowledge of SQL and experience working with structured data.

  • Interest in LLM applications and experience using at least one large language model or AI coding tool.

  • Strong logical thinking and the ability to break an unclear problem into practical, testable steps.

  • A structured approach to experimentation, including tracking changes to prompts, models, data, and parameters and evaluating their impact.

  • Attention to data quality and the willingness to investigate errors rather than treating a successful model response as the end result.

  • Good communication and teamwork skills, with the ability to read English technical documentation.

Nice to Have

  • Project experience involving LLMs, prompt engineering, RAG, agents, information extraction, or automated evaluation.

  • Experience with data processing, ETL, web crawling, data annotation, or knowledge bases.

  • Familiarity with model APIs such as OpenAI, Gemini, or Claude.

  • Exposure to machine learning, NLP, named entity recognition, or structured information extraction.

  • Background in materials science, patents, scientific literature, or technical research data.

  • Experience using AI coding tools such as Cursor, Codex, or Claude Code to build a project, small tool, or automated workflow.

or