⤵️ Quick shortcut: Select the module you’re currently studying ⤵️
Courses Module 1 Module 2 Module 3 Module 4 Module 5 Module 6
By the end of Module 2, you will have developed a comprehensive data pre-processing pipeline for Jarvid, encompassing data cleaning, normalization, tokenization, and efficient handling of large datasets.
Data Cleaning and Normalization:
Use the scripts from Lesson 1 to clean and normalize the employees.csv dataset.
Tokenization:
Implement tokenization as taught in Lesson 2, preparing the data for embedding generation.
Managing Large Datasets:
Optimize the pre-processing pipeline using Dask or PySpark to handle larger volumes of employee data efficiently.
Integration:
Combine all steps into a single, automated Python script that processes raw data and prepares it for the next stage in the pipeline.
A fully functional pre-processing pipeline that transforms raw employee data into a clean, tokenized format ready for embedding and further analysis.
Documentation detailing each step, the tools used, and the reasoning behind chosen methodologies.
To develop a robust and automated data pre-processing pipeline that performs the following tasks:
Cleans and normalizes raw data.
Tokenizes text for embedding generation.
Efficiently handles large datasets for scalability.
Integrates all these steps into a single Python script.
Use scripts from Lesson 1 to:
Remove duplicates: Ensure each row is unique.
Handle missing values: Replace null values with context-specific defaults.
Normalize text: Convert text to lowercase, remove punctuation, and standardize formats.
Code Snippet:
python
Copy code
import pandas as pd
import string
# Load the dataset
raw_file_path = "employees.csv"
df = pd.read_csv(raw_file_path)
# Remove duplicates
df.drop_duplicates(inplace=True)
# Handle missing values
df['role'].fillna('unknown role', inplace=True)
df['department'].fillna('unknown department', inplace=True)
# Normalize text
def clean_text(text):
text = text.lower() # Convert to lowercase
text = text.translate(str.maketrans('', '', string.punctuation)) # Remove punctuation
return text
text_columns = ['name', 'role', 'department']
for column in text_columns:
df[column] = df[column].apply(clean_text)
# Save cleaned data
cleaned_file_path = "employees_cleaned.csv"
df.to_csv(cleaned_file_path, index=False)
print(f"Data cleaning complete. Cleaned file saved to {cleaned_file_path}")
Tokenize the cleaned text using advanced techniques (e.g., spaCy or HuggingFace Tokenizers) for embedding generation.
Code Snippet:
python
Copy code
import spacy
# Load spaCy model
nlp = spacy.load("en_core_web_sm")
# Tokenization function
def tokenize_text(text):
return [token.text for token in nlp(text)]
# Apply tokenization
df['name_tokens'] = df['name'].apply(tokenize_text)
df['role_tokens'] = df['role'].apply(tokenize_text)
df['department_tokens'] = df['department'].apply(tokenize_text)
# Save tokenized data
tokenized_file_path = "employees_tokenized.csv"
df.to_csv(tokenized_file_path, index=False)
print(f"Tokenization complete. Tokenized file saved to {tokenized_file_path}")
Use Dask or PySpark to process larger datasets efficiently.
Option 1: Using Dask
python
Copy code
import dask.dataframe as dd
# Load the dataset using Dask
df = dd.read_csv("employees_tokenized.csv")
# Perform transformations (e.g., filtering or grouping)
filtered_df = df[df['role'].str.contains('engineer')]
grouped_df = filtered_df.groupby('department')['name'].count()
# Compute results
result = grouped_df.compute()
# Save results
optimized_file_path = "employees_optimized.csv"
result.to_csv(optimized_file_path, index=True)
print(f"Optimized data processing complete. Results saved to {optimized_file_path}")
Option 2: Using PySpark
python
Copy code
from pyspark.sql import SparkSession
# Initialize Spark session
spark = SparkSession.builder.appName("Large Dataset Processing").getOrCreate()
# Load dataset into Spark DataFrame
df = spark.read.csv("employees_tokenized.csv", header=True, inferSchema=True)
# Perform transformations
filtered_df = df.filter(df['role'].contains('engineer'))
grouped_df = filtered_df.groupBy("department").count()
# Save results
grouped_df.write.csv("employees_optimized", header=True)
print("Optimized data processing complete.")
Combine all the above steps into a single Python script.
Full Pipeline Script:
python
Copy code
import pandas as pd
import string
import spacy
import dask.dataframe as dd
# Step 1: Data Cleaning
def clean_and_normalize(file_path):
df = pd.read_csv(file_path)
df.drop_duplicates(inplace=True)
df['role'].fillna('unknown role', inplace=True)
df['department'].fillna('unknown department', inplace=True)
def clean_text(text):
text = text.lower()
text = text.translate(str.maketrans('', '', string.punctuation))
return text
for column in ['name', 'role', 'department']:
df[column] = df[column].apply(clean_text)
return df
# Step 2: Tokenization
def tokenize_data(df):
nlp = spacy.load("en_core_web_sm")
def tokenize_text(text):
return [token.text for token in nlp(text)]
df['name_tokens'] = df['name'].apply(tokenize_text)
df['role_tokens'] = df['role'].apply(tokenize_text)
df['department_tokens'] = df['department'].apply(tokenize_text)
return df
# Step 3: Optimized Processing with Dask
def optimize_large_datasets(file_path):
df = dd.read_csv(file_path)
filtered_df = df[df['role'].str.contains('engineer')]
grouped_df = filtered_df.groupby('department')['name'].count()
return grouped_df.compute()
# Main Function
def main():
# File paths
raw_file_path = "employees.csv"
cleaned_file_path = "employees_cleaned.csv"
tokenized_file_path = "employees_tokenized.csv"
optimized_file_path = "employees_optimized.csv"
# Data cleaning
df_cleaned = clean_and_normalize(raw_file_path)
df_cleaned.to_csv(cleaned_file_path, index=False)
# Tokenization
df_tokenized = tokenize_data(df_cleaned)
df_tokenized.to_csv(tokenized_file_path, index=False)
# Large dataset optimization
result = optimize_large_datasets(tokenized_file_path)
result.to_csv(optimized_file_path, index=True)
print("Pre-processing pipeline completed successfully.")
if __name__ == "__main__":
main()
Cleaned Data: employees_cleaned.csv
Tokenized Data: employees_tokenized.csv
Optimized Results: employees_optimized.csv
Include:
Steps performed: Cleaning, tokenization, and large dataset processing.
Tools used: Pandas, spaCy, Dask, or PySpark.
Justification: Why each tool or method was chosen.
Understand the end-to-end data pre-processing pipeline.
Learn to use tools like spaCy, Dask, and PySpark for scalable data handling.
Develop a fully automated script for pre-processing large datasets.