Subject: Ansh Sharma

Code: P-33

Log Date: --.--.----

Observer: Self

APPROVED FOR REVIEW

fingerprint Professional Persona

Subject currently works across data engineering, machine learning, and ML systems. This dossier documents selected projects, experiments, and technical work across data, software, and machine learning, ranging from broader systems work to smaller focused experiments. Individual records cover the problem, implementation, decisions, and results where applicable. The following sections provide a closer look at the work behind the file, including a few unnecessary detours along the way.

Current Status

  • Role:Data Analyst Intern
  • Target:Data Science | ML Engineer
  • Type: Internal Use Only
  • Location: New Delhi, IN
  • Clearance: Level 3

Subject's Whereabouts * linkedin

Service Record
Data Analyst Intern August 2026 - PRESENT
iTech Mission Pvt. Ltd.
  • target Validate government-data modules and trace inconsistencies across source files, frontend inputs, and backend processing.
  • target Prepare and verify master files, indicator metadata, and supporting datasets for ingestion into internal systems.
  • target Perform data cleanup, format checks, and source verification using Excel, SQL, and Python where applicable.
  • target Investigate data-quality issues and recommend validation or preprocessing checks before release.

Experimental Overview (Projects)

menu_book Academia

Data Annotation and Collection Project

Contributed in a academic project on code-mixed Hinglish text classification. Work involved multi-platform data collection across Reddit, Twitter, Instagram, Facebook, YouTube, and Telegram — each with its own access constraints and workarounds.

Also handled corpus annotation and proposed a confidence-threshold based annotation pipeline to reduce manual bottleneck on a ~25k record dataset.

build Technical Core

  • Languages: Python, R, SQL, NoSQL
  • Databases: PostgreSQL, MySQL, MongoDB, Supabase
  • Data: Pandas, NumPy, Apache Arrow, Apache Parquet, DuckDB
  • AI / ML: Scikit-learn, NLTK
  • LLM Tooling: LangChain, Hugging Face Transformers, Ollama, RAG, Vector Embeddings
  • Infrastructure: Docker, Redis, Apache Superset, Flask, Git