MSc Big Data Analytics candidate at Sheffield Hallam University, focused on transforming raw data into reliable analysis, scalable data pipelines and decision-ready business insights.
I build end-to-end analytics projects using Python, SQL, Spark, Hadoop, Hive, machine learning and business intelligence tools. My interests include data analytics, BI and reporting, data engineering, forecasting and responsible applied machine learning.
End-to-end electricity analytics pipeline for Great Britain using official National Energy System Operator data structures.
- Normalises official electricity-demand and system-frequency schemas
- Processes data using Python, PySpark, Hadoop and Hive
- Stores curated outputs as partitioned Parquet tables
- Detects consecutive frequency events outside configured limits
- Compares a Random Forest demand forecast with a lag-1 baseline
- Includes automated tests, an executed notebook and an Excel management report
Technologies: Python · pandas · NumPy · PySpark · Hadoop · Hive · SQL · scikit-learn · R · Jupyter · Excel · pytest · GitHub Actions
Healthcare prescribing analytics project based on NHS Business Services Authority prescribing-data structures.
- Maps official NHSBSA prescribing fields into a curated analytical schema
- Calculates weighted cost-per-item and population-normalised KPIs
- Benchmarks prescribing cost and volume across practices and deprivation groups
- Uses Isolation Forest to identify observations requiring analytical review
- Compares Random Forest forecasting with a previous-month baseline
- Includes Spark, Hive, SQL, testing, documentation and an Excel management report
Technologies: Python · pandas · NumPy · PySpark · Hadoop · Hive · SQL · scikit-learn · Parquet · R · Jupyter · Excel · pytest · GitHub Actions
UK company financial-monitoring and review-prioritisation project using Companies House and XBRL data structures.
- Maps XBRL financial-account fields into company-period analytical records
- Calculates liquidity, liabilities, profitability and net-assets indicators
- Tracks year-on-year financial deterioration
- Benchmarks companies against sector peers
- Produces an explainable rule-based review score
- Separately uses Isolation Forest for statistical review prioritisation
- Includes Spark, Hadoop, Hive, SQL, automated testing and an Excel management report
Technologies: Python · pandas · NumPy · XBRL · PySpark · Hadoop · Hive · SQL · scikit-learn · Parquet · Jupyter · Excel · pytest · GitHub Actions
Customer-retention analytics across 7,043 telecom customer records.
- Performs reproducible data validation and exploratory analysis
- Uses MySQL analytical views and secure database configuration
- Produces customer-segment and churn-driver insights
- Includes an actual Power BI dashboard file
Technologies: Python · pandas · MySQL · SQLAlchemy · SQL · Power BI · data quality · business analysis
Content-based recommendation engine using track-level acoustic features.
- Supports cosine, Euclidean, Manhattan and Pearson similarity
- Applies global feature standardisation
- Generates track and artist recommendations
- Includes command-line and desktop interfaces
- Includes automated tests
Technologies: Python · pandas · NumPy · recommendation systems · similarity metrics · feature engineering · pytest
Local document-analysis application for PDF and TXT files.
- Generates document summaries
- Provides grounded document question answering
- Produces comprehension questions
- Uses a locally hosted Ollama language model
- Includes lightweight document retrieval and automated tests
Technologies: Python · Streamlit · Ollama · LangChain · PDF processing · document retrieval · pytest
| Area | Technologies |
|---|---|
| Data analytics | SQL, Python, R, Excel, SAS |
| Data processing | pandas, NumPy, PySpark, Spark SQL |
| Big data | Hadoop, HDFS, Hive, Parquet |
| Business intelligence | Power BI, Tableau |
| Databases | MySQL, PostgreSQL, MongoDB, SQLite |
| Machine learning | scikit-learn, TensorFlow fundamentals |
| Development | Git, GitHub, Jupyter, Streamlit |
| Quality and automation | pytest, GitHub Actions, data validation |
- Reproducible data cleaning and validation
- SQL analysis and dimensional data modelling
- Scalable Spark, Hadoop and Hive pipelines
- Business-focused KPI design and interpretation
- Forecast evaluation against transparent baselines
- Explainable analytical review methods
- Well-documented and testable repositories
- Honest communication of data and model limitations
- Building a large-scale UK road-accident analytics pipeline using Hadoop, Hive, Spark, R and Tableau/Power BI
- Developing an MSc dissertation on privacy-preserving federated learning for physiological-signal analysis
- Strengthening advanced SQL, dimensional modelling, Power BI/DAX and analytics case-study communication
I am targeting UK graduate and junior opportunities including:
Data Analyst · BI Analyst · MI/Reporting Analyst · Insights Analyst · Junior Data Engineer · Business Analyst · Analytics Graduate Scheme
Explore my repositories for source code, setup instructions, verified outputs, automated tests and project documentation.