Shubham

2D Gaussian Splatting: Geometrically Accurate Radiance Field Reconstruction

Discover how 2D Gaussian Splatting transforms neural rendering by replacing volumetric 3D Gaussians with surface-aligned 2D disks.

3D Computer Graphics, 3D Computer Vision, 3D Reconstruction

VideoRAG: Redefining Long-Context Video Comprehension

Discover VideoRAG, a framework that fuses graph-based reasoning and multi-modal retrieval to enhance LLMs' ability to understand multi-hour videos efficiently.

Agentic AI, LLMs, RAGs, Video Analysis, Vision Language Models

AnomalyCLIP : Harnessing CLIP for Weakly-Supervised Video Anomaly Recognition

Video Anomaly Detection (VAD) is one of the most challenging problems in computer vision. It involves identifying rare, abnormal events in videos – such as burglary, fighting, or accidents –

Anomaly Detection, Vision Transformer, VLMs

Video-RAG: Training-Free Retrieval for Long-Video LVLMs

Learn how Video-RAG boosts training-free and low-compute long-video understanding by pairing OCR, ASR, and open-vocabulary detection with any long-video LVLMs.

RAGs, VLMs

Inside Sinusoidal Position Embeddings: A Sense of Order

In the groundbreaking 2017 paper “Attention Is All You Need”, Vaswani et al. introduced Sinusoidal Position Embeddings to help Transformers encode positional information, without recurrence or convolution. This elegant, non-learned

Language Models, LLMs, NLP

Inside RoPE: Rotary Magic into Position Embeddings

Self-attention, the beating heart of Transformer architectures, treats its input as an unordered set. That mathematical elegance is also a curse: without extra signals, the model has no idea which

Language Models, LLMs, NLP

SmolLM3 Blueprint: SOTA 3B-Parameter LLM

In the evolving landscape of open-source language models, SmolLM3 emerges as a breakthrough: a 3 billion-parameter, decoder-only transformer that rivals larger 4 billion-parameter peers on many benchmarks, while natively supporting

Language Models, LLMs

Fine-Tuning AnomalyCLIP: Class-Agnostic Zero-Shot Anomaly Detection

Zero-shot anomaly detection (ZSAD) is a vital problem in computer vision, particularly in real-world scenarios where labeled anomalies are scarce or unavailable. Traditional vision-language models (VLMs) like CLIP fall short

Anomaly Detection, Vision Transformer, VLMs

Nanonets-OCR-s: Enabling Rich, Structured Markdown for Document Understanding

Traditional Optical Character Recognition (OCR) systems are primarily designed to extract plain text from scanned documents or images. While useful, such systems often ignore semantic structure, layout, and visual cues

OCR, VLMs

Fine-Tuning Grounding DINO: Open-Vocabulary Object Detection

Object detection has traditionally been a closed-set problem: you train on a fixed list of classes and cannot recognize new ones. Grounding DINO breaks this mold, becoming an open-set, language-conditioned