Skip to content
View arijitroy003's full-sized avatar

Organizations

@redhat-data-and-ai

Block or report arijitroy003

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
arijitroy003/README.md

Arijit Kumar Roy — Data and AI systems built for production

Portfolio  ·  LinkedIn  ·  Email

I turn expensive, ambiguous data operations into dependable production systems: agents that shorten incident response, platforms that make governed data self-service, and automation that gives senior engineers their time back.

Today I am a Senior Software Engineer in Data & AI Platform Engineering at Red Hat. Across eight years, I have built for regulated enterprise platforms and consumer products serving up to 120M+ users and 500M events per day.

What I build

01 / Agentic AI

Production MCP and LangChain systems for investigation, release orchestration, metadata intelligence, and operational decision support—with observability and human handoffs designed in.

02 / Data platforms

Self-service data products on Snowflake, Databricks, dbt, Spark, and Delta Lake, with governance, quality, lineage, and cost controls embedded in the paved road.

03 / Platform engineering

Cloud-native control planes and developer workflows built with Python, Go, Kubernetes, OpenShift, GitOps, Terraform, and pragmatic automation.

Systems I have shipped

System What changed Scale and stack
Data Reliability Agent Automated first-line pipeline triage and reduced initial investigation from ~38 minutes to ~4 minutes. MCP, LangChain, OpenShift, Langfuse
Release Assistant Orchestrated governed releases and saved 1,000+ lead-engineer hours annually. Python, GitOps, Snowflake, 150+ compliant data products
Self-service Data Mesh Replaced legacy Redshift/Starburst paths and reduced infrastructure cost by $200K+/year. OpenShift, dbt, Snowflake, Kubernetes
Consumer AI at scale Built GenAI search, recommendations, and conversational systems for Tata Neu and Beem. 120M+ users, 500M events/day, 12 Indic languages

Open source, upstream

I contribute correctness fixes, security hardening, CI improvements, typing support, tests, and documentation across the data and AI ecosystem.

54 merged upstream pull requests across 27 repositories, including the DuckDB, vLLM, Red Hat, dbt Labs, Apache, llm-d, and LangChain communities.

Latest merged upstream pull requests

Selected engineering contributions

Current lab

Small, public experiments where I explore agent interfaces, developer tooling, data workflows, and useful automation.

Project What it explores Language
linkedin-mcp-server LinkedIn automation MCP server wrapping the unofficial linkedin-api Python
snap-a-miro Convert whiteboard photos into interactive Miro boards using AI vision analysis JavaScript
datadiff High-performance CLI tool for semantic diffing of tabular data (CSV, Excel, Parquet, JSON) with Git integration Rust
flight-tracker Local flight price tracker with web UI - supports Amadeus & Skyscanner APIs, daily price monitoring, Indian market optimized Python
Recent public activity

Enterprise contributions

Most of my production work ships to Red Hat's private GitLab. This activity graph provides the missing context that a public GitHub contribution graph cannot.

Red Hat GitLab contribution activity from July 2025 through July 2026

Technical toolkit

languages    Python · Go · Rust · SQL · TypeScript
data         Snowflake · dbt · Databricks · Spark · Delta Lake · Kafka · Airflow
ai systems   LangChain · MCP · OpenAI · Claude · Mistral · Vector DBs · Langfuse
platform     Kubernetes · OpenShift · GitOps · Terraform · Docker · AWS · Azure · GCP

MCA, Jadavpur University · Distributed systems and information retrieval research at ISI Kolkata


Building a serious data platform or production AI system?
Start a conversation  ·  Explore my work

Project, activity, and upstream contribution data refresh automatically through GitHub Actions.

Pinned Loading

  1. arijitroy003.github.io arijitroy003.github.io Public

    Personal portfolio website showcasing my work in Data Engineering, AI/ML, and Software Development

    JavaScript

  2. datadiff datadiff Public

    High-performance CLI tool for semantic diffing of tabular data (CSV, Excel, Parquet, JSON) with Git integration

    Rust

  3. linkedin-mcp-server linkedin-mcp-server Public

    LinkedIn automation MCP server wrapping the unofficial linkedin-api

    Python

  4. snap-a-miro snap-a-miro Public

    Convert whiteboard photos into interactive Miro boards using AI vision analysis

    JavaScript

  5. duckdb duckdb Public

    Forked from duckdb/duckdb

    DuckDB is an analytical in-process SQL database management system

    C++

  6. genai-toolbox genai-toolbox Public

    Forked from googleapis/mcp-toolbox

    MCP Toolbox for Databases is an open source MCP server for databases.

    Go