Hi, I'm Aritro Roy
- 22-year-old B.Tech Computer Science graduate from Asansol, India.
- Software Developer with hands-on experience in backend engineering, data engineering, and AI/ML.
- I am active on X. Follow me on X!
- Actively learning Rust, AI systems and System Design
- Outside of coding, I enjoy gaming, performing music live, and weightlifting.
- I'm passionate about building architectures that can scale easily, automating workflows, and drinking a lot of caffeine.
- Always curious, always learning.
GitHub Activity
953 contributions in the last year
Skills
Languages
Frameworks & Tools
Data Engineering & AI
Databases
Cloud & DevOps
Experience
Software Development Engineer
Whatbytes Technologies | Remote
- Architected a multi-tenant document intelligence platform for commercial real-estate leases on FastAPI — webhook-driven sync across Dropbox, Box, Google Drive and SharePoint (Nango, GCS) feeding a Vertex AI OCR and extraction pipeline that structures 100+ key points per lease into isolated Supabase schemas.
- Cut end-to-end ingestion latency 60% by parallelizing Celery async stages and optimising schema throughput.
- Designed Airbyte CDK ETL connectors (full and incremental syncs, 20+ SaaS entities) and a Django REST backend for an AI red-teaming platform, with zero-code YAML tool registration and dependency-graph onboarding.
- Uncovered and independently verified the production fix for an OAuth defect blocking Google Drive integration in a Rust (Axum, Tokio, SQLite) governance daemon with a Svelte dashboard.
- Engineered a Stripe-backed billing engine in Django REST on manual-capture PaymentIntents, atomic rollback and row-level locking (select_for_update), eliminating double-spend and duplicate-redemption races across every purchase flow.
FastAPIVertex AICeleryAirbyte CDKRustStripeOctober, 2025 -Present
Backend Engineering Intern
UrbanRider Technologies | Telangana, Hyderabad (Remote)
- Resolved PostgreSQL query bottlenecks and REST API inefficiencies to cut application load latency 22%, restructuring indexing to hold performance under peak traffic.
- Removed manual release steps by automating validation, regression testing and CI/CD with Python, Docker and GitHub Actions, ending environment-parity bugs between staging and production.
PostgreSQLDockerGitHub ActionsJuly, 2022 -May, 2023
Side Projects
A live voice-to-answer RAG service over a 99K-passage corpus (ElevenLabs STT -> FAISS -> cross-encoder rerank -> Claude) on AWS EC2, holding a P50 of 96ms against a 200ms budget. Cross-encoder reranking lifted recall@5 from 0.848 to 0.916, with per-sentence groundedness guards catching 100% of ungrounded answers at zero false refusals.
A FastAPI query router dispatching requests across Qdrant vector and structured data stores, with tool-use agentic workflows and async concurrency for multi-user workloads, improving semantic query accuracy 25% over a keyword baseline.
An AI-powered cold email platform for automated outreach campaigns, with contact extraction, personalized templating, scheduling, and a response-tracking dashboard.
a cloud-native resume analysis tool with real-time PDF parsing and LLM integration.
A Python scraper for extracting structured case data from public court records, built for reliable, repeatable data collection at scale.
Impact
60%
faster document ingestion after parallelizing Celery stages
96ms
P50 latency for the RAG pipeline, against a 200ms budget
25%
higher semantic query accuracy over a keyword baseline
0.848 -> 0.916
recall@5 after adding cross-encoder reranking
22%
cut in application load latency from index restructuring