Content
<div align="center">
<h1>
<img src="https://readme-typing-svg.demolab.com?font=JetBrains+Mono&weight=800&size=42&duration=3000&pause=1000&color=6E40C9¢er=true&vCenter=true&width=800&lines=DataMinds+Labs;AI-Powered+Dataset+Intelligence;From+Raw+Ideas+%E2%86%92+ML-Ready+Data" alt="DataMinds Labs" />
</h1>
<br/>
<p align="center">
<a href="https://github.com/dhruv-005/DataMind_Labs">
<img src="https://img.shields.io/badge/⚡_DataMinds_Labs-v1.0.0-6E40C9?style=for-the-badge&labelColor=0D1117" />
</a>
</p>
<p align="center">
<a href="https://github.com/dhruv-005/DataMind_Labs/stargazers">
<img src="https://img.shields.io/github/stars/dhruv-005/DataMind_Labs?style=for-the-badge&logo=github&color=f0c040&labelColor=0D1117&logoColor=white" />
</a>
<a href="https://github.com/dhruv-005/DataMind_Labs/network/members">
<img src="https://img.shields.io/github/forks/dhruv-005/DataMind_Labs?style=for-the-badge&logo=github&color=4A90D9&labelColor=0D1117&logoColor=white" />
</a>
<a href="https://github.com/dhruv-005/DataMind_Labs/issues">
<img src="https://img.shields.io/github/issues/dhruv-005/DataMind_Labs?style=for-the-badge&logo=github&color=E05C5C&labelColor=0D1117&logoColor=white" />
</a>
<a href="https://github.com/dhruv-005/DataMind_Labs/blob/main/LICENSE">
<img src="https://img.shields.io/badge/License-MIT-22C55E?style=for-the-badge&labelColor=0D1117&logo=opensourceinitiative&logoColor=white" />
</a>
<a href="https://github.com/dhruv-005/DataMind_Labs/commits/main">
<img src="https://img.shields.io/github/last-commit/dhruv-005/DataMind_Labs?style=for-the-badge&logo=git&color=F97316&labelColor=0D1117&logoColor=white" />
</a>
</p>
<p align="center">
<img src="https://img.shields.io/badge/Python-3.10%2B-3776AB?style=for-the-badge&logo=python&logoColor=white&labelColor=0D1117" />
<img src="https://img.shields.io/badge/Ollama-LLaMA3%20%7C%20Mistral-FF6B35?style=for-the-badge&logo=ollama&logoColor=white&labelColor=0D1117" />
<img src="https://img.shields.io/badge/MCP-Tool%20Protocol-8B5CF6?style=for-the-badge&logoColor=white&labelColor=0D1117" />
<img src="https://img.shields.io/badge/UI-Streamlit-FF4B4B?style=for-the-badge&logo=streamlit&logoColor=white&labelColor=0D1117" />
<img src="https://img.shields.io/badge/Status-Active%20Development-22C55E?style=for-the-badge&logoColor=white&labelColor=0D1117" />
</p>
<br/>
<p align="center">
<b>Autonomous · Intelligent · End-to-End · Production-Grade</b>
</p>
<p align="center">
The open-source AI platform that transforms raw ideas into<br/>
bias-audited, ML-ready datasets — without writing a single line of preprocessing code.
</p>
<br/>
```
┌─────────────────────────────────────────────────────────────────────────┐
│ Generate → Clean → Annotate → Process → Audit → Report │
│ One prompt. Six intelligent stages. │
└─────────────────────────────────────────────────────────────────────────┘
```
<br/>
</div>
---
## Table of Contents
<details>
<summary><b>Expand Full Navigation</b></summary>
<br/>
| # | Section |
|---|---------|
| 01 | [Overview](#-overview) |
| 02 | [Core Mission](#-core-mission) |
| 03 | [System Architecture](#-system-architecture) |
| 04 | [Module Breakdown](#-module-breakdown) |
| 05 | [Module 1 — AI Dataset Generator](#-module-1--ai-dataset-generator) |
| 06 | [Module 2 — Data Cleaner](#-module-2--data-cleaner) |
| 07 | [Module 3 — Data Annotator & Labeler](#-module-3--data-annotator--labeler) |
| 08 | [Module 4 — Data Processor](#-module-4--data-processor) |
| 09 | [Module 5 — Quality Auditor](#-module-5--quality-auditor) |
| 10 | [Module 6 — Report Generator](#-module-6--report-generator) |
| 11 | [Module 7 — MCP Tool Layer](#-module-7--mcp-tool-layer) |
| 12 | [Project Structure](#-project-structure) |
| 13 | [Tech Stack](#-tech-stack) |
| 14 | [Installation & Setup](#-installation--setup) |
| 15 | [Usage Guide](#-usage-guide) |
| 16 | [Sample Output](#-sample-output) |
| 17 | [Development Roadmap](#-development-roadmap) |
| 18 | [Contributing](#-contributing) |
| 19 | [License](#-license) |
| 20 | [Author](#-author) |
</details>
---
<br/>
## Overview
<div align="center">
> **"Stop spending 80% of your time on data preparation. Let the pipeline do it for you."**
</div>
**DataMinds Labs** is a production-grade, autonomous AI dataset intelligence platform engineered for machine learning engineers, data scientists, researchers, and developers who need high-quality datasets — fast, clean, and audit-verified.
This is not a data cleaning wrapper. This is a **complete, intelligent data pipeline system** — from domain-specific synthetic data generation all the way through bias auditing and professional PDF reporting — orchestrated by an autonomous agent layer built on the **Model Context Protocol (MCP).**
<br/>
### Why DataMinds Labs?
| Problem | DataMinds Labs Solution |
|---------|------------------------|
| Dataset collection is time-consuming and costly | Generate domain-specific synthetic datasets in minutes using LLMs |
| Manual data cleaning is error-prone and repetitive | Automated multi-strategy cleaning with full audit trail |
| Labeling requires expensive human annotation | Zero-shot LLM annotation with ML-verified confidence scoring |
| Feature engineering requires deep domain knowledge | Auto feature engineering with intelligent type detection |
| Bias and leakage go undetected until production | Built-in fairness auditor, leakage detector, and diversity scorer |
| No unified visibility across the pipeline | AI-generated PDF reports with visualizations and recommendations |
<br/>
---
## Core Mission
```
╔══════════════════════════════════════════════════════════════════════════╗
║ ║
║ To democratize high-quality dataset creation by making the entire ║
║ data pipeline — from generation to audit — accessible, automated, ║
║ and intelligent for every developer, researcher, and data ║
║ professional — regardless of infrastructure or budget. ║
║ ║
╚══════════════════════════════════════════════════════════════════════════╝
```
<br/>
## System Architecture
The entire DataMinds Labs pipeline is orchestrated by a central LLM agent that routes tasks across six intelligent modules through the MCP tool protocol.
```
╔══════════════════════════════════════════════════════════════════════════════╗
║ DATAMINDS LABS — SYSTEM PIPELINE ║
╠══════════════════════════════════════════════════════════════════════════════╣
║ ║
║ INPUT ║
║ ┌──────────────────────────────────────────────────────────────────────┐ ║
║ │ User Prompt: "Generate a fraud detection dataset with 10,000 rows" │ ║
║ └─────────────────────────────────┬────────────────────────────────────┘ ║
║ │ ║
║ ▼ ║
║ ┌──────────────────────────────────────────────────────────────────────┐ ║
║ │ LLM AGENT (Orchestrator) │ ║
║ │ Ollama · LLaMA 3 · Model Context Protocol │ ║
║ └──────┬───────────┬───────────┬───────────┬───────────┬───────────────┘ ║
║ │ │ │ │ │ ║
║ ▼ ▼ ▼ ▼ ▼ ║
║ ┌────────────┐ ┌──────────┐ ┌──────────┐ ┌─────────┐ ┌──────────────┐ ║
║ │ MODULE 1 │ │ MODULE 2 │ │ MODULE 3 │ │MODULE 4 │ │ MODULE 5 │ ║
║ │ Generator │ │ Cleaner │ │Annotator │ │Processor│ │ Auditor │ ║
║ └──────┬─────┘ └────┬─────┘ └────┬─────┘ └────┬────┘ └──────┬───────┘ ║
║ │ │ │ │ │ ║
║ └────────────┴────────────┴────────────┴─────────────┘ ║
║ │ ║
║ ▼ ║
║ ┌──────────────────────────────────────────────────────────────────────┐ ║
║ │ MODULE 6 — REPORT GENERATOR │ ║
║ │ PDF Report · Visualizations · AI Recommendations │ ║
║ └─────────────────────────────────┬────────────────────────────────────┘ ║
║ │ ║
║ ▼ ║
║ OUTPUT ║
║ ┌──────────────────────────────────────────────────────────────────────┐ ║
║ │ CSV · JSON · Excel · PDF Report · Dashboard │ ║
║ └──────────────────────────────────────────────────────────────────────┘ ║
║ ║
╚══════════════════════════════════════════════════════════════════════════════╝
```
<br/>
---
## Module Breakdown
<div align="center">
| Module | Name | Core Responsibility | Key Technologies |
|:------:|:-----|:-------------------|:----------------|
| `M1` | AI Dataset Generator | Generate domain-specific synthetic datasets from plain language prompts | Ollama, Faker, SDV |
| `M2` | Data Cleaner | Multi-strategy cleaning: missing values, duplicates, outliers, noise, type errors | Pandas, Scikit-learn |
| `M3` | Data Annotator & Labeler | Zero-shot LLM annotation with ML-verified confidence scoring | Ollama, SpaCy, NLTK |
| `M4` | Data Processor | Feature engineering, encoding, scaling, class balancing, train/test splits | Scikit-learn, SMOTE |
| `M5` | Quality Auditor | Completeness, consistency, bias, leakage, diversity, and ML readiness scoring | SHAP, Pandas, SciPy |
| `M6` | Report Generator | Professional PDF reports with charts, metrics, and AI-generated recommendations | ReportLab, Plotly |
| `M7` | MCP Tool Layer | Agent-to-tool routing and orchestration via Model Context Protocol | MCP SDK, FastAPI |
</div>
<br/>
---
## Module 1 — AI Dataset Generator
The generator module accepts plain language prompts and produces structured, diverse, domain-accurate synthetic datasets using an LLM-powered schema inference engine.
```
┌─────────────────────────────────────────────────────────────────────┐
│ AI DATASET GENERATOR — FLOW │
├─────────────────────────────────────────────────────────────────────┤
│ │
│ User Prompt │
│ │ │
│ ▼ │
│ LLM Domain Parser │
│ Understands domain, identifies entities, infers column types │
│ │ │
│ ▼ │
│ Schema Builder │
│ Column names · Data types · Constraints · Relationships │
│ │ │
│ ▼ │
│ Row-by-Row Structured Generation │
│ LLM + Faker + SDV generate realistic, non-repetitive rows │
│ │ │
│ ▼ │
│ Diversity Validator │
│ Checks distribution, entropy, and value range coverage │
│ │ │
│ ▼ │
│ Structured Dataset Output │
│ CSV · JSON · Excel · SQLite │
│ │
└─────────────────────────────────────────────────────────────────────┘
```
<br/>
### Supported Dataset Domains
| Domain | Dataset Examples | Max Rows |
|--------|-----------------|:--------:|
| NLP & Text | Sentiment Analysis, Spam Detection, Intent Classification, Toxicity | 50,000 |
| Healthcare | Symptom-Diagnosis, Patient Records, Drug Response, ICU Readmission | 50,000 |
| Finance | Fraud Detection, Credit Scoring, Transaction History, Loan Approval | 50,000 |
| E-Commerce | Product Reviews, Sales Forecasting, Customer Segmentation, Churn | 50,000 |
| HR & People Analytics | Employee Attrition, Performance, Burnout Prediction, Engagement | 50,000 |
| Education | Student Performance, Dropout Risk, Learning Outcome Prediction | 50,000 |
| Cybersecurity | Network Intrusion Logs, Attack Classification, Anomaly Detection | 50,000 |
| Custom Domain | Any domain specified in plain English — fully dynamic schema | 50,000 |
<br/>
### Prompt Examples
```bash
# NLP Dataset
"Generate a sentiment analysis dataset for restaurant reviews — 5,000 rows"
# Healthcare Dataset
"Generate a medical symptom-to-diagnosis dataset with 10,000 patient records"
# Finance Dataset
"Generate a credit card fraud detection dataset with realistic transaction patterns"
# HR Dataset
"Generate an employee burnout prediction dataset with mental health and productivity indicators"
# Custom Domain
"Generate a social media post engagement prediction dataset with platform, post type, and reach metrics"
```
## Module 2 — Data Cleaner
The cleaner module applies a comprehensive multi-strategy pipeline to every uploaded or generated dataset. Each strategy is configurable and tracked for the final audit report.
```
┌─────────────────────────────────────────────────────────────────────┐
│ DATA CLEANER — PIPELINE │
├─────────────────────────────────────────────────────────────────────┤
│ │
│ Raw Dataset Input │
│ │ │
│ ├── MISSING VALUE HANDLER │
│ │ · Mean / Median / Mode Imputation │
│ │ · KNN Imputation (for correlated columns) │
│ │ · Predictive ML Imputation (IterativeImputer) │
│ │ · Auto-drop columns exceeding 50% missingness │
│ │ │
│ ├── DUPLICATE DETECTOR │
│ │ · Exact duplicate row removal │
│ │ · Near-duplicate fuzzy matching (token similarity) │
│ │ │
│ ├── OUTLIER HANDLER │
│ │ · IQR-based detection │
│ │ · Z-Score statistical detection │
│ │ · Isolation Forest (unsupervised ML) │
│ │ · User-configurable action: Keep · Remove · Cap │
│ │ │
│ ├── NOISE REDUCER │
│ │ · Spelling correction and normalization │
│ │ · Format standardization (dates, phone numbers) │
│ │ · Whitespace trimming and case unification │
│ │ │
│ ├── DATA TYPE ENFORCER │
│ │ · Auto type inference from column content │
│ │ · Date format standardization (ISO 8601) │
│ │ · Numeric string-to-float conversion │
│ │ │
│ ├── INCONSISTENCY RESOLVER │
│ │ · Value canonicalization │
│ │ │
│ └── NLP TEXT CLEANER │
│ · HTML and XML tag removal │
│ · Special character and emoji stripping │
│ · Stopword removal (optional, configurable) │
│ · Contraction expansion │
│ │
│ Clean Dataset Output → /data/cleaned/ │
│ │
└─────────────────────────────────────────────────────────────────────┘
```
<br/>
---
## Module 3 — Data Annotator & Labeler
The annotator module implements a two-stage labeling pipeline — LLM zero-shot classification followed by ML model verification — to produce high-confidence, audit-traceable labels at scale.
```
┌─────────────────────────────────────────────────────────────────────┐
│ AUTO-LABELING PIPELINE — ARCHITECTURE │
├─────────────────────────────────────────────────────────────────────┤
│ │
│ Unlabeled Data Input │
│ │ │
│ ▼ │
│ Stage 1 — LLM Zero-Shot Annotation │
│ Classifies text or numeric records using contextual reasoning │
│ │ │
│ ▼ │
│ Stage 2 — ML Model Verification │
│ Cross-validates labels using a trained classifier │
│ │ │
│ ▼ │
│ Stage 3 — Confidence Score Assignment │
│ Per-label probability scored from agreement between both stages │
│ │ │
│ ├── Confidence ≥ 85% → Auto-Approved │
│ └── Confidence < 85% → Flagged for Review │
│ │
│ Labeled Dataset Output → /data/labeled/ │
│ │
└─────────────────────────────────────────────────────────────────────┘
```
<br/>
### Annotation Types
| Category | Supported Tasks |
|----------|----------------|
| **Text Annotation** | Sentiment (Positive / Negative / Neutral), Intent Detection, Named Entity Recognition, Topic Classification, Toxicity Detection, Emotion Detection (6-class) |
| **Numeric Annotation** | Risk Level (High / Medium / Low), Anomaly Flagging, Cluster Assignment, Percentile Banding |
<br/>
### Confidence Score Output Format
```
┌─────────────────────────────────────────────────────────────────┐
│ ANNOTATION RESULT SAMPLE │
├─────────────────────────────────────────────────────────────────┤
│ │
│ Record : "The food was absolutely outstanding." │
│ Label : POSITIVE │
│ Score : ████████████████████ 96% │
│ Method : Full Agreement │
│ │
├─────────────────────────────────────────────────────────────────┤
│ │
│ Record : "Service was kind of okay I guess." │
│ Label : NEUTRAL │
│ Score : ██████████████░░░░░░ 71% │
│ Method : Partial Agreement — Flagged for Review │
│ │
└─────────────────────────────────────────────────────────────────┘
```
<br/>
---
## Module 4 — Data Processor
The processor module prepares clean, labeled data for direct consumption by machine learning training pipelines. Every transformation is deterministic, reproducible, and logged.
```
┌─────────────────────────────────────────────────────────────────────┐
│ DATA PROCESSOR — PIPELINE │
├─────────────────────────────────────────────────────────────────────┤
│ │
│ Cleaned & Labeled Data Input │
│ │ │
│ ├── FEATURE ENGINEERING │
│ │ · Date decomposition → Year, Month, Day, Quarter │
│ │ · Text features → Word count, Char count, Sentiment │
│ │ · Numeric derivations → Ratios, Differences, Log │
│ │ · Interaction terms for correlated columns │
│ │ │
│ ├── CATEGORICAL ENCODING │
│ │ · Label Encoding (binary & ordinal columns) │
│ │ · One-Hot Encoding (nominal low-cardinality columns) │
│ │ · Target Encoding (high-cardinality columns) │
│ │ · Ordinal Encoding (auto-detected rank order) │
│ │ │
│ ├── NORMALIZATION & SCALING │
│ │ · Min-Max Scaling (bounded distributions) │
│ │ · Standard Scaling / Z-Score (Gaussian features) │
│ │ · Robust Scaling (outlier-resistant) │
│ │ · Log Transformation (right-skewed distributions) │
│ │ │
│ ├── CLASS IMBALANCE HANDLER │
│ │ · SMOTE Oversampling (synthetic minority generation) │
│ │ · Random Undersampling │
│ │ · Class Weight Adjustment │
│ │ · Hybrid SMOTE + Undersampling │
│ │ │
│ ├── TRAIN / VALIDATION / TEST SPLIT │
│ │ · Stratified splitting (preserves class ratios) │
│ │ · Configurable split ratios (default 80/10/10) │
│ │ · K-Fold Cross Validation support │
│ │ │
│ ├── DIMENSIONALITY REDUCTION │
│ │ · PCA (configurable explained variance threshold) │
│ │ · Correlation-based feature elimination │
│ │ · Variance Threshold filtering │
│ │ │
│ └── TEXT VECTORIZATION │
│ · TF-IDF (term frequency weighting) │
│ · Word Embeddings (Word2Vec / GloVe) │
│ · Sentence Transformers (HuggingFace) │
│ │
│ ML-Ready Dataset Output → /data/processed/ │
│ │
└─────────────────────────────────────────────────────────────────────┘
```
<br/>
---
## Module 5 — Quality Auditor
The quality auditor runs a structured, multi-dimensional health assessment on every dataset. It produces a quantified score across six dimensions and surfaces actionable warnings automatically.
```
┌─────────────────────────────────────────────────────────────────────┐
│ QUALITY AUDITOR — PIPELINE │
├─────────────────────────────────────────────────────────────────────┤
│ │
│ Processed Dataset Input │
│ │ │
│ ├── COMPLETENESS AUDIT │
│ │ · Missing percentage per column and globally │
│ │ · Required column coverage check │
│ │ │
│ ├── CONSISTENCY AUDIT │
│ │ · Value format uniformity per column │
│ │ · Cross-column logical rule validation │
│ │ │
│ ├── ACCURACY AUDIT │
│ │ · Outlier density per numeric feature │
│ │ · Noise level analysis and quantification │
│ │ │
│ ├── BIAS DETECTION │
│ │ · Class imbalance ratio per target column │
│ │ · Demographic representation analysis │
│ │ (gender, age group, geographic distribution) │
│ │ · SHAP-based feature influence fairness check │
│ │ │
│ ├── DATA LEAKAGE DETECTION │
│ │ · Target-feature correlation sweep │
│ │ · Temporal leakage detection (future data bleed) │
│ │ · ID and proxy feature identification │
│ │ │
│ ├── DIVERSITY SCORE │
│ │ · Shannon entropy per categorical feature │
│ │ · Distribution spread across numeric ranges │
│ │ · Unique value coverage ratio │
│ │ │
│ └── ML READINESS SCORE │
│ · Composite score across all six dimensions │
│ · Recommended model families per dataset type │
│ · Export readiness assessment │
│ │
└─────────────────────────────────────────────────────────────────────┘
```
<br/>
---
## Module 6 — Report Generator
The report generator produces a professional, self-contained PDF audit report for every processed dataset. Reports include visualizations, scoring breakdowns, and AI-authored recommendations.
<br/>
### Report Contents
| Section | Description |
|---------|-------------|
| **Dataset Overview** | Row count, column summary, data types, file size, generation timestamp |
| **Generation Details** | Original prompt, detected domain, inferred schema, generation duration |
| **Cleaning Summary** | Rows affected, strategies applied, values imputed, columns modified |
| **Annotation Summary** | Labels applied, confidence distribution histogram, flagged record count |
| **Quality Scores** | All six audit dimension scores with visual progress indicators |
| **Visualizations** | Distribution plots, correlation heatmap, missing value map, class balance chart, feature importance chart |
| **AI Recommendations** | Automatically generated, issue-specific suggestions for every warning raised |
| **Warnings & Risk Flags** | Critical issues that require human review before model training |
### Sample AI Recommendation Block
```
┌─────────────────────────────────────────────────────────────────┐
│ AI RECOMMENDATIONS — EXCERPT │
├─────────────────────────────────────────────────────────────────┤
│ │
│ ⚠ WARNING: Column 'age' — 12% missing values detected. │
│ Strategy Applied: KNN Imputation using correlated features │
│ │
│ ⚠ WARNING: Class imbalance detected in target column. │
│ Distribution → Class 0: 85% · Class 1: 15% │
│ Strategy Applied: SMOTE Oversampling (ratio 1:1) │
│ │
│ ✓ SUGGESTION: Apply Log1p transformation to 'income' │
│ column — right skew detected (skewness = 2.34) │
│ │
│ ✓ SUGGESTION: Drop 'employee_id' — zero variance, │
│ no predictive signal, potential leakage risk │
│ │
│ ✓ MODEL RECOMMENDATION: │
│ XGBoost or LightGBM — optimal for this dataset based │
│ on feature types, class distribution, and cardinality │
│ │
│ ✓ DIVERSITY: Shannon entropy score = HIGH across all │
│ categorical features. Excellent training coverage. │
│ │
└─────────────────────────────────────────────────────────────────┘
```
<br/>
---
## Module 7 — MCP Tool Layer
The MCP (Model Context Protocol) tool layer is the agent orchestration backbone of DataMinds Labs. The central LLM agent dynamically selects and calls MCP-registered tools based on task context.
```
┌─────────────────────────────────────────────────────────────────────┐
│ MCP AGENT — ARCHITECTURE │
├─────────────────────────────────────────────────────────────────────┤
│ │
│ ┌────────────────────────────────────────┐ │
│ │ LLM AGENT │ │
│ │ Ollama · LLaMA 3 · Orchestrator │ │
│ └──────────────────┬─────────────────────┘ │
│ │ │
│ ▼ │
│ ┌────────────────────────────────────────┐ │
│ │ MCP TOOL ROUTER │ │
│ │ Parses intent → Selects tool(s) │ │
│ └──┬──────┬──────┬──────┬──────┬─────────┘ │
│ │ │ │ │ │ │
│ ▼ ▼ ▼ ▼ ▼ │
│ ┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐ │
│ │ File │ │ Gen │ │Clean │ │Annot │ │Audit │ │
│ │ Tool │ │ Tool │ │ Tool │ │ Tool │ │ Tool │ │
│ └──────┘ └──────┘ └──────┘ └──────┘ └──────┘ │
│ │ │
│ ┌────────────────┘ │
│ ▼ │
│ ┌────────────────────────────────────────┐ │
│ │ REPORT + EXPORT TOOL │ │
│ │ PDF · CSV · JSON · Excel │ │
│ └────────────────────────────────────────┘ │
│ │
└─────────────────────────────────────────────────────────────────────┘
```
<br/>
### MCP Server Registry
| Server File | Exposed Tools | Responsibility |
|-------------|--------------|----------------|
| `file_server.py` | `read_file`, `write_file`, `list_files` | Dataset file I/O operations |
| `generator_server.py` | `generate_dataset`, `build_schema`, `validate_diversity` | AI-powered data generation |
| `cleaner_server.py` | `clean_data`, `detect_outliers`, `fix_types`, `remove_duplicates` | Data quality cleaning |
| `annotator_server.py` | `annotate_text`, `label_numeric`, `score_confidence` | LLM + ML annotation |
| `auditor_server.py` | `audit_quality`, `detect_bias`, `check_leakage`, `score_health` | Multi-dimension auditing |
| `report_server.py` | `generate_report`, `export_pdf`, `create_charts`, `write_recommendations` | PDF report generation |
<br/>
---
## Project Structure
```
DataMinds_Labs/
│
├── mcp_servers/ # MCP Tool Layer — Server definitions
│ ├── file_server.py # File system read/write operations
│ ├── generator_server.py # Dataset generation MCP tools
│ ├── cleaner_server.py # Data cleaning MCP tools
│ ├── annotator_server.py # Annotation and labeling MCP tools
│ ├── auditor_server.py # Quality audit MCP tools
│ └── report_server.py # Report generation MCP tools
│
├── modules/ # Core module implementations
│ │
│ ├── generator/ # Module 1 — AI Dataset Generator
│ │ ├── dataset_generator.py # Main LLM-powered generation engine
│ │ ├── schema_builder.py # Domain schema inference and construction
│ │ └── diversity_checker.py # Output diversity validation and entropy check
│ │
│ ├── cleaner/ # Module 2 — Data Cleaner
│ │ ├── missing_handler.py # Multi-strategy missing value imputation
│ │ ├── outlier_detector.py # IQR, Z-Score, and Isolation Forest detection
│ │ ├── duplicate_remover.py # Exact and fuzzy duplicate elimination
│ │ ├── noise_reducer.py # Format standardization and normalization
│ │ └── text_cleaner.py # NLP-specific text preprocessing
│ │
│ ├── annotator/ # Module 3 — Data Annotator & Labeler
│ │ ├── text_annotator.py # Sentiment, NER, intent, and topic labeling
│ │ ├── numeric_labeler.py # Risk, anomaly, and cluster label assignment
│ │ └── confidence_scorer.py # Per-label confidence computation
│ │
│ ├── processor/ # Module 4 — Data Processor
│ │ ├── feature_engineer.py # Auto feature creation and derivation
│ │ ├── encoder.py # Categorical encoding strategy engine
│ │ ├── scaler.py # Normalization and scaling pipeline
│ │ ├── imbalance_handler.py # SMOTE and resampling strategies
│ │ └── splitter.py # Stratified train/validation/test splitting
│ │
│ ├── auditor/ # Module 5 — Quality Auditor
│ │ ├── completeness_checker.py # Missing value audit and coverage scoring
│ │ ├── bias_detector.py # Fairness analysis and demographic audit
│ │ ├── leakage_detector.py # Target correlation and temporal leakage check
│ │ ├── diversity_scorer.py # Entropy and distribution diversity analysis
│ │ └── health_scorer.py # Composite health score computation
│ │
│ └── reporter/ # Module 6 — Report Generator
│ ├── report_generator.py # Full report orchestration logic
│ ├── visualizer.py # Charts, plots, heatmaps, and dashboards
│ └── pdf_exporter.py # Professional PDF generation via ReportLab
│
├── agent/ # LLM Agent and Tool Orchestration
│ ├── main_agent.py # Central orchestrating LLM agent
│ └── tool_router.py # MCP tool intent classification and routing
│
├── ui/ # Streamlit Web Interface
│ └── app.py # Full interactive pipeline dashboard
│
├── data/ # Dataset storage (gitignored)
│ ├── raw/ # Original generated or uploaded datasets
│ ├── cleaned/ # Post-cleaning output datasets
│ ├── labeled/ # Annotated and labeled datasets
│ └── processed/ # ML-ready final datasets
│
├── reports/ # Audit report output
│ └── output_reports/ # Generated PDF and HTML reports
│
├── requirements.txt # Python dependency manifest
├── .env.example # Environment variable template
├── .gitignore # Git ignore configuration
└── README.md # Project documentation
```
<br/>
---
## Tech Stack
<div align="center">
| Category | Technology | Version | Purpose |
|----------|-----------|:-------:|---------|
| **LLM Engine** | Ollama (LLaMA 3 / Mistral) | Latest | AI generation, zero-shot annotation, orchestration |
| **Agent Protocol** | MCP Python SDK | Latest | Tool-calling architecture and agent routing |
| **Backend API** | FastAPI | 0.110+ | REST API layer for MCP server communication |
| **Core Language** | Python | 3.10+ | All module implementations |
| **Data Processing** | Pandas + NumPy | Latest | Data manipulation, transformation, and I/O |
| **ML Layer** | Scikit-learn + XGBoost | Latest | Cleaning support, annotation verification, processing |
| **Synthetic Data** | Faker + SDV | Latest | Realistic and statistically valid data generation |
| **Class Balancing** | Imbalanced-learn | Latest | SMOTE and resampling strategies |
| **NLP Toolkit** | SpaCy + NLTK + HuggingFace | Latest | Text processing, NER, and sentence embeddings |
| **Visualization** | Matplotlib + Seaborn + Plotly | Latest | Charts, heatmaps, and interactive dashboards |
| **PDF Generation** | ReportLab | Latest | Professional PDF export for audit reports |
| **UI Dashboard** | Streamlit | Latest | Interactive web interface for the full pipeline |
| **Database** | SQLite | Built-in | Lightweight local dataset and session storage |
| **Explainability** | SHAP | Latest | Feature importance analysis for bias detection |
</div>
<br/>
---
## Installation & Setup
### System Requirements
```
Python >= 3.10
RAM >= 8 GB (16 GB recommended for datasets > 10K rows)
Storage >= 5 GB free disk space
Ollama >= Latest (required for LLM-powered features)
OS : macOS · Linux · Windows (WSL2 recommended)
```
<br/>
### Step 1 — Clone the Repository
```bash
git clone https://github.com/dhruv-005/DataMind_Labs.git
cd DataMind_Labs
```
<br/>
### Step 2 — Create and Activate Virtual Environment
```bash
# Create the virtual environment
python -m venv venv
# Activate on macOS / Linux
source venv/bin/activate
# Activate on Windows
venv\Scripts\activate
```
<br/>
### Step 3 — Install All Dependencies
```bash
pip install --upgrade pip
pip install -r requirements.txt
```
<br/>
### Step 4 — Install and Configure Ollama
```bash
# Install Ollama on macOS / Linux
curl -fsSL https://ollama.ai/install.sh | sh
# Pull the primary model (LLaMA 3)
ollama pull llama3
# Pull optional alternative model (Mistral)
ollama pull mistral
# Verify installation
ollama list
```
<br/>
### Step 5 — Configure Environment Variables
```bash
# Copy the example environment file
cp .env.example .env
# Edit with your configuration
nano .env
```
```env
# ─── Ollama Configuration ───────────────────────────────────────────────────
OLLAMA_BASE_URL=http://localhost:11434
OLLAMA_MODEL=llama3
# ─── Data Directory Paths ───────────────────────────────────────────────────
RAW_DATA_DIR=./data/raw
CLEANED_DATA_DIR=./data/cleaned
LABELED_DATA_DIR=./data/labeled
PROCESSED_DATA_DIR=./data/processed
REPORTS_DIR=./reports/output_reports
# ─── Pipeline Defaults ──────────────────────────────────────────────────────
DEFAULT_CONFIDENCE_THRESHOLD=0.85
DEFAULT_MISSING_VALUE_THRESHOLD=0.50
DEFAULT_TRAIN_SPLIT=0.80
DEFAULT_VALIDATION_SPLIT=0.10
DEFAULT_TEST_SPLIT=0.10
DEFAULT_RANDOM_STATE=42
```
<br/>
### Step 6 — Verify Installation
```bash
python -c "
import pandas, numpy, sklearn, faker, spacy, streamlit, reportlab, plotly
print('All core dependencies verified successfully.')
"
```
<br/>
---
## Usage Guide
### Option A — Streamlit Dashboard (Recommended)
Launch the full interactive web interface:
```bash
streamlit run ui/app.py
```
Navigate to `http://localhost:8501` in your browser. The dashboard provides step-by-step access to every module, with live previews and downloadable outputs at each stage.
<br/>
### Option B — CLI Agent Mode
Run the orchestrating agent directly from the terminal:
```bash
# Interactive mode
python agent/main_agent.py
# Inline prompt
python agent/main_agent.py --prompt "Generate a fraud detection dataset with 5,000 rows"
# Full options
python agent/main_agent.py \
--prompt "Generate sentiment analysis dataset for restaurant reviews" \
--rows 10000 \
--format csv \
--output ./data/raw/restaurant_reviews.csv
```
<br/>
### Option C — Full Pipeline — Single Command
Execute the complete end-to-end pipeline:
```bash
python agent/main_agent.py \
--prompt "Generate employee attrition prediction dataset" \
--rows 8000 \
--run-full-pipeline \
--export-report \
--output-format csv
```
<br/>
### Option D — Module-Level Python API
Use individual modules directly in your own code:
```python
```
# Generate a dataset programmatically
from modules.generator.dataset_generator import DatasetGenerator
generator = DatasetGenerator(model="llama3")
dataset = generator.generate(
prompt="Credit card fraud detection dataset with transaction metadata",
rows=5000,
output_format="csv"
)
print(f"Generated {len(dataset)} rows successfully.")
```python
# Run a quality audit on an existing dataset
from modules.auditor.health_scorer import HealthScorer
import pandas as pd
df = pd.read_csv("data/cleaned/my_dataset.csv")
scorer = HealthScorer()
report = scorer.audit(df, target_column="label")
print(report.summary())
```
```python
# Clean an existing dataset
from modules.cleaner.missing_handler import MissingHandler
from modules.cleaner.outlier_detector import OutlierDetector
import pandas as pd
df = pd.read_csv("data/raw/raw_data.csv")
df = MissingHandler(strategy="knn").fit_transform(df)
df = OutlierDetector(method="isolation_forest").fit_transform(df)
df.to_csv("data/cleaned/clean_data.csv", index=False)
```
<br/>
---
## Sample Output
### Dataset Health Report
```
╔══════════════════════════════════════════════════════════════════════╗
║ DATAMINDS LABS — DATASET AUDIT REPORT ║
║ Generated: 2025-01-15 · Version: 1.0.0 ║
╠══════════════════════════════════════════════════════════════════════╣
║ Dataset : employee_attrition_dataset.csv ║
║ Rows : 8,000 · Columns : 24 ║
║ Domain : HR & People Analytics ║
╠══════════════════════════════════════════════════════════════════════╣
║ ║
║ AUDIT DIMENSION SCORES ║
║ ──────────────────────────────────────────────────────────────── ║
║ Completeness ████████████████████░░ 94% EXCELLENT ✓ ║
║ Consistency ████████████████░░░░░░ 88% GOOD ✓ ║
║ Accuracy █████████████████████░ 91% EXCELLENT ✓ ║
║ Bias Score ████████████████████░░ LOW ✓ ║
║ Diversity █████████████████████░ HIGH ✓ ║
║ ML Readiness ████████████████░░░░░░ 89% GOOD ✓ ║
║ ║
╠══════════════════════════════════════════════════════════════════════╣
║ ║
║ OVERALL HEALTH SCORE ║
║ ║
║ ████████████████████████████████████░░░░ 87 / 100 ║
║ ║
║ GRADE : A- [Excellent] ║
║ ║
╠══════════════════════════════════════════════════════════════════════╣
║ ║
║ WARNINGS ║
║ ⚠ Column 'age' — 12% missing → KNN Imputation applied ║
║ ⚠ Class imbalance: 0 = 79% · 1 = 21% → SMOTE applied ║
║ ║
║ SUGGESTIONS ║
║ ✓ Apply Log1p transform to 'monthly_income' (skew = 2.34) ║
║ ✓ Drop 'employee_id' — zero predictive signal, leakage risk ║
║ ║
║ RECOMMENDED MODELS ║
║ [1] XGBoost [2] LightGBM [3] Random Forest ║
╚══════════════════════════════════════════════════════════════════════╝
```
<br/>
---
## Development Roadmap
```
╔══════════════════════════════════════════════════════════════════════╗
║ DEVELOPMENT ROADMAP — 6 WEEKS ║
╠══════════╦═══════════════════════════════════════════════════════════╣
║ ║ ║
║ WEEK 1 ║ Project Foundation & MCP Architecture ║
║ ║ · Environment setup, dependency audit, folder structure ║
║ ║ · All six MCP servers scaffolded and interconnected ║
║ ║ · Ollama integration verified end-to-end ║
║ ║ · Git repository initialized with CI checks ║
║ ║ ║
╠══════════╬═══════════════════════════════════════════════════════════╣
║ ║ ║
║ WEEK 2 ║ AI Dataset Generator Module ║
║ ║ · LLM domain parser and schema builder ║
║ ║ · Row-by-row structured generation engine ║
║ ║ · Diversity checker with entropy validation ║
║ ║ · CSV, JSON, Excel, and SQLite export support ║
║ ║ ║
╠══════════╬═══════════════════════════════════════════════════════════╣
║ ║ ║
║ WEEK 3 ║ Data Cleaner & Processor Modules ║
║ ║ · Full cleaning pipeline (missing, outliers, noise) ║
║ ║ · Feature engineering, encoding, and scaling engine ║
║ ║ · SMOTE imbalance handler and stratified splitter ║
║ ║ ║
╠══════════╬═══════════════════════════════════════════════════════════╣
║ ║ ║
║ WEEK 4 ║ Data Annotator & Labeler Module ║
║ ║ · LLM zero-shot annotation pipeline ║
║ ║ · ML verification layer for label validation ║
║ ║ · Confidence scoring with auto-approve thresholds ║
║ ║ ║
╠══════════╬═══════════════════════════════════════════════════════════╣
║ ║ ║
║ WEEK 5 ║ Quality Auditor Module ║
║ ║ · All six audit dimensions implemented ║
║ ║ · SHAP-based bias and feature influence detection ║
║ ║ · Composite health score calculator ║
║ ║ ║
╠══════════╬═══════════════════════════════════════════════════════════╣
║ ║ ║
║ WEEK 6 ║ Report Generator + Streamlit UI + Integration Testing ║
║ ║ · PDF report generation with full visualization suite ║
║ ║ · Streamlit dashboard — live interactive interface ║
║ ║ · End-to-end pipeline integration and regression tests ║
║ ║ · README finalized and repository published ║
║ ║ ║
╚══════════╩═══════════════════════════════════════════════════════════╝
```
<br/>
---
## Contributing
Contributions are what make open-source projects genuinely useful. All contributions to DataMinds Labs are warmly welcomed — from bug fixes and documentation improvements to entirely new module capabilities.
### How to Contribute
```bash
# 1. Fork the repository on GitHub
# 2. Clone your fork
git clone https://github.com/YOUR_USERNAME/DataMind_Labs.git
cd DataMind_Labs
# 3. Create a descriptive feature branch
git checkout -b feature/your-feature-name
# 4. Make your changes and commit with a clear message
git add .
git commit -m "feat: describe what you added or changed"
# 5. Push to your fork
git push origin feature/your-feature-name
# 6. Open a Pull Request on GitHub with a clear description
```
<br/>
### Contribution Areas
| Area | What We Need |
|------|-------------|
| **New Dataset Domains** | Domain schemas, example prompts, and generation templates |
| **Cleaning Strategies** | New imputation methods, noise reduction techniques |
| **Audit Checks** | Additional quality dimensions, fairness metrics |
| **UI Improvements** | Streamlit component enhancements, UX improvements |
| **Bug Reports** | Detailed issues with reproduction steps and environment info |
| **Documentation** | Module-level docstrings, usage examples, tutorials |
<br/>
### Commit Message Convention
```
feat: New feature or capability added
fix: Bug fix or error correction
docs: Documentation addition or update
style: Code formatting with no logic change
refactor: Code restructure with no behavior change
test: Test additions or updates
chore: Build configuration, dependency, or tooling update
```
<br/>
---
## License
```
MIT License
Copyright (c) 2026 Dhruv Sonani
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in
all copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN
THE SOFTWARE.
```
See the full [`LICENSE`](./LICENSE) file for details.
<br/>
---
## Author
<div align="center">
<br/>
### Dhruv Sonani
*AI & Machine Learning Engineer · Data Pipeline Architect · Open Source Builder*
<br/>
[](https://github.com/dhruv-005)
[](https://linkedin.com/in/dhruv-sonani)
[](mailto:your-email@example.com)
[](https://github.com/dhruv-005/DataMind_Labs)
<br/>
> *"Building DataMinds Labs to prove that great data infrastructure should be*
> *open, intelligent, and accessible to every builder on the planet."*
<br/>
---
### Support the Project
If DataMinds Labs helped you, saved you time, or inspired your own work —
consider giving the repository a star. It genuinely helps the project reach more developers.
<br/>
[](https://github.com/dhruv-005/DataMind_Labs)
<br/>
---
<p align="center">
<sub>
DataMinds Labs — Built with precision by
<a href="https://github.com/dhruv-005">Dhruv Sonani</a>
· MIT Licensed · Open Source Forever
</sub>
</p>
</div>
Connection Info
You Might Also Like
markitdown
MarkItDown-MCP is a lightweight server for converting URIs to Markdown.
markitdown
Python tool for converting files and office documents to Markdown.
Filesystem
Node.js MCP Server for filesystem operations with dynamic access control.
TrendRadar
TrendRadar: Your hotspot assistant for real news in just 30 seconds.
mempalace
The highest-scoring AI memory system ever benchmarked. And it's free.
mempalace
The highest-scoring AI memory system ever benchmarked. And it's free.