Skip to content
View govardhanreddi's full-sized avatar

Block or report govardhanreddi

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
govardhanreddi/README.md

Hi, I'm Govardhan Reddy

MSc Big Data Analytics candidate at Sheffield Hallam University, focused on transforming raw data into reliable analysis, scalable data pipelines and decision-ready business insights.

I build end-to-end analytics projects using Python, SQL, Spark, Hadoop, Hive, machine learning and business intelligence tools. My interests include data analytics, BI and reporting, data engineering, forecasting and responsible applied machine learning.

Featured Analytics Projects

End-to-end electricity analytics pipeline for Great Britain using official National Energy System Operator data structures.

  • Normalises official electricity-demand and system-frequency schemas
  • Processes data using Python, PySpark, Hadoop and Hive
  • Stores curated outputs as partitioned Parquet tables
  • Detects consecutive frequency events outside configured limits
  • Compares a Random Forest demand forecast with a lag-1 baseline
  • Includes automated tests, an executed notebook and an Excel management report

Technologies: Python · pandas · NumPy · PySpark · Hadoop · Hive · SQL · scikit-learn · R · Jupyter · Excel · pytest · GitHub Actions


Healthcare prescribing analytics project based on NHS Business Services Authority prescribing-data structures.

  • Maps official NHSBSA prescribing fields into a curated analytical schema
  • Calculates weighted cost-per-item and population-normalised KPIs
  • Benchmarks prescribing cost and volume across practices and deprivation groups
  • Uses Isolation Forest to identify observations requiring analytical review
  • Compares Random Forest forecasting with a previous-month baseline
  • Includes Spark, Hive, SQL, testing, documentation and an Excel management report

Technologies: Python · pandas · NumPy · PySpark · Hadoop · Hive · SQL · scikit-learn · Parquet · R · Jupyter · Excel · pytest · GitHub Actions


UK company financial-monitoring and review-prioritisation project using Companies House and XBRL data structures.

  • Maps XBRL financial-account fields into company-period analytical records
  • Calculates liquidity, liabilities, profitability and net-assets indicators
  • Tracks year-on-year financial deterioration
  • Benchmarks companies against sector peers
  • Produces an explainable rule-based review score
  • Separately uses Isolation Forest for statistical review prioritisation
  • Includes Spark, Hadoop, Hive, SQL, automated testing and an Excel management report

Technologies: Python · pandas · NumPy · XBRL · PySpark · Hadoop · Hive · SQL · scikit-learn · Parquet · Jupyter · Excel · pytest · GitHub Actions


Customer-retention analytics across 7,043 telecom customer records.

  • Performs reproducible data validation and exploratory analysis
  • Uses MySQL analytical views and secure database configuration
  • Produces customer-segment and churn-driver insights
  • Includes an actual Power BI dashboard file

Technologies: Python · pandas · MySQL · SQLAlchemy · SQL · Power BI · data quality · business analysis


Content-based recommendation engine using track-level acoustic features.

  • Supports cosine, Euclidean, Manhattan and Pearson similarity
  • Applies global feature standardisation
  • Generates track and artist recommendations
  • Includes command-line and desktop interfaces
  • Includes automated tests

Technologies: Python · pandas · NumPy · recommendation systems · similarity metrics · feature engineering · pytest


Local document-analysis application for PDF and TXT files.

  • Generates document summaries
  • Provides grounded document question answering
  • Produces comprehension questions
  • Uses a locally hosted Ollama language model
  • Includes lightweight document retrieval and automated tests

Technologies: Python · Streamlit · Ollama · LangChain · PDF processing · document retrieval · pytest

Technical Toolkit

Area Technologies
Data analytics SQL, Python, R, Excel, SAS
Data processing pandas, NumPy, PySpark, Spark SQL
Big data Hadoop, HDFS, Hive, Parquet
Business intelligence Power BI, Tableau
Databases MySQL, PostgreSQL, MongoDB, SQLite
Machine learning scikit-learn, TensorFlow fundamentals
Development Git, GitHub, Jupyter, Streamlit
Quality and automation pytest, GitHub Actions, data validation

What I Focus On

  • Reproducible data cleaning and validation
  • SQL analysis and dimensional data modelling
  • Scalable Spark, Hadoop and Hive pipelines
  • Business-focused KPI design and interpretation
  • Forecast evaluation against transparent baselines
  • Explainable analytical review methods
  • Well-documented and testable repositories
  • Honest communication of data and model limitations

Current Work

  • Building a large-scale UK road-accident analytics pipeline using Hadoop, Hive, Spark, R and Tableau/Power BI
  • Developing an MSc dissertation on privacy-preserving federated learning for physiological-signal analysis
  • Strengthening advanced SQL, dimensional modelling, Power BI/DAX and analytics case-study communication

Career Interests

I am targeting UK graduate and junior opportunities including:

Data Analyst · BI Analyst · MI/Reporting Analyst · Insights Analyst · Junior Data Engineer · Business Analyst · Analytics Graduate Scheme


Explore my repositories for source code, setup instructions, verified outputs, automated tests and project documentation.

Popular repositories Loading

  1. genai-document-assistant genai-document-assistant Public

    Local GenAI document assistant for PDF summarisation, grounded Q&A and question generation.

    Python 1

  2. telecom-churn-analytics telecom-churn-analytics Public

    End-to-end customer churn analytics using Python, SQL, MySQL and Power BI.

    Python 1

  3. Music-Recommendation-and-Analysis-System Music-Recommendation-and-Analysis-System Public

    Content-based music recommendation system using Python and multiple similarity algorithms.

    Python 1

  4. gb-gridpulse gb-gridpulse Public

    Big-data analytics for GB electricity demand, renewable generation and frequency stability.

    Jupyter Notebook 1

  5. nhs-rxinsight nhs-rxinsight Public

    Healthcare prescribing analytics using Python, Spark, SQL, Hive and population-normalised NHS metrics.

    Jupyter Notebook 1

  6. uk-businesswatch uk-businesswatch Public

    UK company financial analytics using Python, XBRL, Spark, Hive and explainable review scoring.

    Python 1