Adam Woods

Data Scientist - Machine Learning & Generative AI

I work with machine learning and generative AI to find the best solutions to real-world problems, building knowledge and skills as I go.

10 +Projects
6 YearsIn Data Science
LocationBased in Leeds, UK

About

01

I like learning and solving new problems.

Most of my work is focused around research and experimentation for problems that require rigorous solutions, with the findings I create actionable insights to help make every decision informed and effective, or when possible, create robust, long-term solutions.

As long as I'm learning, I'm happy.

Location
Leeds, UK - open to remote working
Focus
Research, experimentation, tools and everything in between
Tools
Python, SQL, Git, PyTorch, LangChain and many more
Currently
Senior Data Scientist at PwC
Education
1st Degree in Comp Sci with AI

Expertise

02

Machine Learning

Forecasting, classification and causal inference on real, messy operational data.

Neural Networksscikit-learnXGBoostPyTorchFeature Extraction

Generative AI

LLM red-teaming and evaluation, retrieval-augmented systems, and prompt optimisation.

LangChainJudgeLLMDSPyToken AnalysisHugging Face

Research & Experimentation

Finding and testing the best methods for solving real-world challenges in machine learning and quantum computing.

Research PapersKaggleLiterature ReviewDocumentation

Data Engineering

Developing pipelines, tools and products to automate machine learning systems for repeated use with dynamic data.

AutomationDeliveryReportingVersion Management

Experience

04
???? - ????

Insert Your Company Here?

Working hard and doing cool stuff!

Who Knows
2024 - Now

Senior Data Scientist, PwC

At PwC my work was focused on academic research, leading experimentation projects, gathering insight and delivering pipelines to implement the latest techniques. My focus was around three domains, AI Safety, Quantum Computing and Prompt Optimisation, which I was leading for the research team. This role included management of projects and junior colleagues

Leeds
2021 - 2024

Senior Data Scientist, UK Search Limited

As the only data scientist at the company my role covered a wide variety of topics, from data engineering to machine learning. I also worked with two others in the strategy team, guiding them through SQL and Python. The work I did was heavily centred around clients with and operational strategy.

Sheffield

Projects

03
2023

Payment Classification

A machine learning classification project using a neural network trained on over 10,000 data points to determine whether debt collection customers were likely to make a payment, enabling more informed decisions about which customers to contact and helping to significantly reduce costs.

Machine LearningNeural NetworkClassificationAnomaly DetectionCost Reduction
2025

LLM Red-teaming Evaluation

A production ready pipeline to determine if an LLM response contained dangerous or unwanted behaviour. Included LLM Judge feature extraction combined with an XGBoost model to align with human evaluation. This is now being used to evaluate real agentic tools built by clients.

Machine LearningGenerative AILLM JudgeFeature Extraction
2026

Quantum Computing Research

A long-term, Knowledge gathering project into the rapidly developing area of quantum computing. The aim is to develop a team of subject matter experts that can help the firm answer any questions clients may have about quantum computing, a growing trend that will be very impactful in the coming years. I learned the basics of quantum physics and also experimented with quantum algorithm simulation.

Quantum ComputingAlgorithmsDocumentation
2026

Prompt Optimisation

For the AI Research team I led the prompt optimisation domain. This focused on applying optimisation techniques to LLM prompts in order to improve task performance and discover best practices for agentic system design. This led to the development of a prompt optimisation pipeline, a dynamic tool that allows us to easily optimise prompts for any scenario and test a wide range of methods to ensure confident results.

OptimisationAgentic SystemsPrompt EngineeringProject Management
2026

Token Optimisation

An experimental project looking into token optimisation with the aim to help reduce costs for LLM systems, a growing concern for internal teams and clients. This work involved research and testing into several token reduction and prompt compression techniques. Our findings helped many teams within the firm find solutions suitable for their use case.

Cost ReductionPrompt EngineeringOptimisationProject Management
2020

Exo-Planet Identification

My dissertation project using data from NASA to identify extremely distant stars with planets orbiting them based on recorded light intensity over-time data. At the time this type of work was done mostly by subject matter experts manually, taking significant time. This project used a Convolutional Neural Network to identify exo-planets with an accuracy over 80%.

Neural NetworkMachine LearningAstronomyTime Series Data
2024

Token Safeguarding

An AI safety project around LLM blue-teaming with an aim of preventing jailbreak attacks. Several methods were tested but my part focused on token analysis, using metrics such as attention and perplexity to identify tokens that could lead to harmful responses and remove them before the model generates a response. These findings helped to identify the requirements and use cases for these techniques as well as the limitations such as context collapse.

AI SafetyGenerative AIBlue-teamingToken Analysis
2026

Agentic Red-Teaming

This AI safety project was focused around agentic systems, an especially important topic as LLMs now have the ability to take actions that have real world consequences. Testing these models with agentic specific techniques helped identify vulnerabilities that must be considered when building agentic systems and highlighted models that were reliable for agentic use cases.

AI SafetyGenerative AIRed-teamingAgentic Systems
2021

Report Automation

For UKSL one major service was to provide performance updates to clients on a regular basis, with some clients even demanding daily reports. Before I joined the team this was done manually! To solve this, I built a pipeline that allowed us to create a simple YAML configuration file to set up reports to be scheduled and sent fully autonomously. This implementation meant non-technical users could set up reports easily and freed up staff to work on more interesting areas.

AutomationStakeholder ManagementPandasMatplotlib
2025

Code Red-Teaming

Another LLM safety project focused on testing models for malicious code generation. This involved implementing code specific jailbreak techniques such as pseudo-code adaptation, separating malicious requests into multiple benign components to be combined later and providing an empty function for the model to fill out to force the priority onto fulfilling the request over safety concerns. This project revealed the vulnerabilities in models and highlighted models that were especially susceptible to such attacks.

AI SafetyGenerative AIRed-TeamingPrompt Engineering
2025

LLM Data Exfiltration

A project focused on red-teaming retrieval augmented generation systems to identify model vulnerabilities that could impact data privacy. This involved research and experimentation for several techniques such as data poisoning with redirection and prompt injection attacks. This research demonstrated the importance of security regarding data accessible to LLMs and how content filtering should be used to avoid any sensitive data leakage.

AI SafetyGenerative AIRAGData Poisoning
2021

Light Curve Search

As part of my degree, this project built upon my existing work with NASA's data on the light intensity of a star over time. This project used fuzzy logic and Fréchet distance to match a user drawing of a light curve to one from a real star, then linking the user to the NASA archive for that star. This project was a bit of fun but was great for learning to handle time-series data.

AstronomyTime Series DataFuzzy Logic
2026

Small Language Model Safety

With the growing popularity of locally hosted small language models to save costs and give the organisation power over data security. The concern over safety testing was significant. This project gathered extensive research into the landscape of SLM red-teaming and found that today's SLMs are not built with safety in mind and should be used with caution and only for extremely narrow tasks. It also highlighted the vulnerabilities that come with locally hosting models, shifting the power but also the responsibility away from the large AI labs and onto the hosting organisation.

AI SafetyGenerative AIFine-TuningSecurity

Knowledge

05

Below is a network graph of my personal notes, it has been filtered for display purposes so only shows around 17% of my notes. It's still a bit messy but hopefully demonstrates the scale of my knowledge and my organisational skills when it comes to learning. You can hover or tap to see a preview of what the note covers.

Hub page Page Tag

Drag the canvas to pan, scroll or pinch to zoom. The note summaries are AI generated so might not be perfectly accurate to the original notes. All my notes are designed to be general knowledge and contain very little -sensitive data. Any -sensitive data has been removed from these notes.

Heard enough?

The fastest way to reach me is by email, or find me on LinkedIn below.

adam@adamwoods.dev

If you do contact me, please let me know if you looked at my website, it would be interesting to find out who looks at this stuff.