Mastering NLP from Foundations to LLMs

By : Lior Gazit, Meysam Ghaffari

Mastering NLP from Foundations to LLMs

By: Lior Gazit, Meysam Ghaffari

Overview of this book

Do you want to master Natural Language Processing (NLP) but don’t know where to begin? This book will give you the right head start. Written by leaders in machine learning and NLP, Mastering NLP from Foundations to LLMs provides an in-depth introduction to techniques. Starting with the mathematical foundations of machine learning (ML), you’ll gradually progress to advanced NLP applications such as large language models (LLMs) and AI applications. You’ll get to grips with linear algebra, optimization, probability, and statistics, which are essential for understanding and implementing machine learning and NLP algorithms. You’ll also explore general machine learning techniques and find out how they relate to NLP. Next, you’ll learn how to preprocess text data, explore methods for cleaning and preparing text for analysis, and understand how to do text classification. You’ll get all of this and more along with complete Python code samples. By the end of the book, the advanced topics of LLMs’ theory, design, and applications will be discussed along with the future trends in NLP, which will feature expert opinions. You’ll also get to strengthen your practical skills by working on sample real-world NLP business problems and solutions.

Preface

Who this book is for

What this book covers

To get the most out of this book

Download the example code files

Conventions used

Get in touch

Reviews

Share Your Thoughts

Download a free PDF copy of this book

Free Chapter

Chapter 1: Navigating the NLP Landscape: A Comprehensive Introduction

Who this book is for

What is natural language processing?

Initial strategies in the machine processing of natural language

A winning synergy – the coming together of NLP and ML

Introduction to math and statistics in NLP

Summary

Questions and answers

Chapter 2: Mastering Linear Algebra, Probability, and Statistics for Machine Learning and NLP

Introduction to linear algebra

Eigenvalues and eigenvectors

Basic probability for machine learning

Summary

Further reading

References

Chapter 3: Unleashing Machine Learning Potentials in Natural Language Processing

Technical requirements

Data exploration

Common machine learning models

Model underfitting and overfitting

Splitting data

Hyperparameter tuning

Ensemble models

Handling imbalanced data

Dealing with correlated data

Summary

References

Chapter 4: Streamlining Text Preprocessing Techniques for Optimal NLP Performance

Technical requirements

Lowercasing in NLP

Removing special characters and punctuation

NER

POS tagging

Explaining the preprocessing pipeline

Summary

Chapter 5: Empowering Text Classification: Leveraging Traditional Machine Learning Techniques

Technical requirements

Types of text classification

Text classification using TF-IDF

Text classification using Word2Vec

Topic modeling – a particular use case of unsupervised text classification

Reviewing our use case – ML system design for NLP classification in a Jupyter Notebook

Summary

Chapter 6: Text Classification Reimagined: Delving Deep into Deep Learning Language Models

Technical requirements

Understanding deep learning basics

The architecture of different neural networks

The challenges of training neural networks

Language models

Understanding transformers

Learning more about large language models

The challenges of training language models

Challenges of using GPT-3

Summary

Chapter 7: Demystifying Large Language Models: Theory, Design, and Langchain Implementation

Technical requirements

What are LLMs and how are they different from LMs?

How LLMs stand out

Motivations for developing and using LLMs

Challenges in developing LLMs

Different types of LLMs

Example designs of state-of-the-art LLMs

Summary

References

Chapter 8: Accessing the Power of Large Language Models: Advanced Setup and Integration with RAG

Technical requirements

Setting up an LLM application – API-based closed source models

Prompt engineering and priming GPT

Setting up an LLM application – local open source models

Employing LLMs from Hugging Face via Python

Exploring advanced system design – RAG and LangChain

Reviewing a simple LangChain setup in a Jupyter notebook

LLMs in the cloud

Summary

Chapter 9: Exploring the Frontiers: Advanced Applications and Innovations Driven by LLMs

Technical requirements

Enhancing LLM performance with RAG and LangChain – a dive into advanced functionalities

Advanced methods with chains

Retrieving information from various web sources automatically

Prompt compression and API cost reduction

Multiple agents – forming a team of LLMs that collaborate

Summary

Chapter 10: Riding the Wave: Analyzing Past, Present, and Future Trends Shaped by LLMs and AI

Key technical trends around LLMs and AI

Large datasets and their indelible mark on NLP and LLMs

Evolution of large language models – purpose, value, and impact

NLP and LLMs in the business world

Behavioral trends induced by AI and LLMs – the social aspect

Summary

Chapter 11: Exclusive Industry Insights: Perspectives and Predictions from World Class Experts

Overview of our experts

Nitzan Mekel-Bobrov, PhD

David Sontag, PhD

John D. Halamka, M.D., M.S.

Xavier Amatriain, PhD

Melanie Garson, PhD

Our questions and the experts’ answers

Summary

Index

Why subscribe?

Other Books You May Enjoy

Packt is searching for authors like you

Share Your Thoughts

Download a free PDF copy of this book

Customer Reviews

5 star

4 star

3 star

2 star

1 star

What is natural language processing?

NLP is a field of artificial intelligence (AI) focused on the interaction between computers and human languages. It involves using computational techniques to understand, interpret, and generate human language, making it possible for computers to understand and respond to human input naturally and meaningfully.

The history and evolution of natural language processing

The history of NLP is a fascinating journey through time, tracing back to the 1950s, with significant contributions from pioneers such as Alan Turing. Turing’s seminal paper, Computing Machinery and Intelligence, introduced the Turing test, laying the groundwork for future explorations in AI and NLP. This period marked the inception of symbolic NLP, characterized by the use of rule-based systems, such as the notable Georgetown experiment in 1954, which ambitiously aimed to solve machine translation by generating a translation of Russian content into English (see https://en.wikipedia.org/wiki/Georgetown%E2%80%93IBM_experiment). Despite early optimism, progress was slow, revealing the complexities of language understanding and generation.

The 1960s and 1970s saw the development of early NLP systems, which demonstrated the potential for machines to engage in human-like interactions using limited vocabularies and knowledge bases. This era also witnessed the creation of conceptual ontologies, crucial for structuring real-world information in a computer-understandable format. However, the limitations of rule-based methods led to a paradigm shift in the late 1980s towards statistical NLP, fueled by advances in ML and increased computational power. This shift enabled more effective learning from large corpora, significantly advancing machine translation and other NLP tasks. This paradigm shift not only represented a technological and methodological advancement but also underscored a conceptual evolution in the approach to linguistics within NLP. In moving away from the rigidity of predefined grammar rules, this transition embraced corpus linguistics, a method that allows machines to “perceive” and understand languages through extensive exposure to large bodies of text. This approach reflects a more empirical and data-driven understanding of language, where patterns and meanings are derived from actual language use rather than theoretical constructs, enabling more nuanced and flexible language processing capabilities.

Entering the 21st century, the emergence of the web provided vast amounts of data, catalyzing research in unsupervised and semi-supervised learning algorithms. The breakthrough came with the advent of neural NLP in the 2010s, where DL techniques began to dominate, offering unprecedented accuracy in language modeling and parsing. This era has been marked by the development of sophisticated models such as Word2Vec and the proliferation of deep neural networks, driving NLP towards more natural and effective human-computer interaction. As we continue to build on these advancements, NLP stands at the forefront of AI research, with its history reflecting a relentless pursuit of understanding and replicating the nuances of human language.

In recent years, NLP has also been applied to a wide range of industries, such as healthcare, finance, and social media, where it has been used to automate decision-making and enhance communication between humans and machines. For example, NLP has been used to extract information from medical documents, analyze customer feedback, translate documents between languages, and search through enormous amounts of posts.

Mastering NLP from Foundations to LLMs

By : Lior Gazit, Meysam Ghaffari

Mastering NLP from Foundations to LLMs

By: Lior Gazit, Meysam Ghaffari

Overview of this book

Related Content you might be interested in

Current Title:

Mastering NLP from Foundations to LLMs

ChatGPT Prompts Book - Precision Prompts, Priming, Training & AI Writing Techniques for Mortals

GPT-3

Generative AI for Cloud Solutions

What is natural language processing?

The history and evolution of natural language processing