<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Vishal Rajpurohit — Articles</title><description>In-depth technical writing on AI systems, software architecture, Laravel, and rescuing software that is in trouble.</description><link>https://vishal.life/</link><language>en</language><copyright>© 2026 Vishal Rajpurohit</copyright><item><title>Transformer Inference, From HTTP Request to the Next Token</title><link>https://vishal.life/blog/transformer-inference-from-request-to-token/</link><guid isPermaLink="true">https://vishal.life/blog/transformer-inference-from-request-to-token/</guid><description>A systems-level tour of prefill, decode, batching, memory movement, and the scheduling decisions that determine real-world LLM latency.</description><pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate><category>Inference Systems</category><category>transformers</category><category>inference</category><category>GPU</category><category>serving</category><author>Vishal Rajpurohit</author></item><item><title>KV Cache Engineering: Pages, Prefix Reuse, and Memory Pressure</title><link>https://vishal.life/blog/kv-cache-engineering-paged-attention-prefix-reuse/</link><guid isPermaLink="true">https://vishal.life/blog/kv-cache-engineering-paged-attention-prefix-reuse/</guid><description>How to size, allocate, reuse, compress, and evict the state that makes autoregressive decoding practical.</description><pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate><category>Inference Systems</category><category>KV cache</category><category>paged attention</category><category>memory</category><category>latency</category><author>Vishal Rajpurohit</author></item><item><title>Semantic Caching Without Serving the Wrong Answer Faster</title><link>https://vishal.life/blog/semantic-caching-with-correctness-boundaries/</link><guid isPermaLink="true">https://vishal.life/blog/semantic-caching-with-correctness-boundaries/</guid><description>A production design for similarity keys, freshness, authorization, invalidation, and the economics that determine when semantic reuse is safe.</description><pubDate>Sun, 12 Jul 2026 00:00:00 GMT</pubDate><category>Inference Systems</category><category>semantic cache</category><category>embeddings</category><category>latency</category><category>correctness</category><author>Vishal Rajpurohit</author></item><item><title>Speculative Decoding Without Hand-Waving</title><link>https://vishal.life/blog/speculative-decoding-production-guide/</link><guid isPermaLink="true">https://vishal.life/blog/speculative-decoding-production-guide/</guid><description>Draft models, acceptance math, tree proposals, and the operational details that decide whether speculation makes serving faster.</description><pubDate>Fri, 10 Jul 2026 00:00:00 GMT</pubDate><category>Inference Systems</category><category>speculative decoding</category><category>latency</category><category>draft models</category><category>sampling</category><author>Vishal Rajpurohit</author></item><item><title>Synthetic Evaluation Data That Finds Real Failures</title><link>https://vishal.life/blog/synthetic-data-for-evaluations/</link><guid isPermaLink="true">https://vishal.life/blog/synthetic-data-for-evaluations/</guid><description>How to generate, filter, diversify, and maintain synthetic cases without turning an evaluation suite into a mirror of its generator.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate><category>Evaluation and Security</category><category>synthetic data</category><category>evaluation</category><category>coverage</category><category>LLM judges</category><author>Vishal Rajpurohit</author></item><item><title>The LLM Gateway Is a Policy Engine, Not a Proxy</title><link>https://vishal.life/blog/llm-gateway-as-policy-and-control-plane/</link><guid isPermaLink="true">https://vishal.life/blog/llm-gateway-as-policy-and-control-plane/</guid><description>Designing the routing, budgets, identity, resilience, and evidence layer between products and a changing model portfolio.</description><pubDate>Mon, 06 Jul 2026 00:00:00 GMT</pubDate><category>Inference Systems</category><category>LLM gateway</category><category>model routing</category><category>governance</category><category>resilience</category><author>Vishal Rajpurohit</author></item><item><title>Privacy-Preserving AI Is a Dataflow Architecture</title><link>https://vishal.life/blog/privacy-preserving-production-ai/</link><guid isPermaLink="true">https://vishal.life/blog/privacy-preserving-production-ai/</guid><description>Purpose limitation, minimization, isolation, retention, redaction, and verifiable deletion across retrieval, models, tools, traces, and feedback loops.</description><pubDate>Fri, 03 Jul 2026 00:00:00 GMT</pubDate><category>Evaluation and Security</category><category>privacy</category><category>data governance</category><category>PII</category><category>security</category><author>Vishal Rajpurohit</author></item><item><title>Incident Response for Systems That Can Be Wrong Fluently</title><link>https://vishal.life/blog/incident-response-for-ai-systems/</link><guid isPermaLink="true">https://vishal.life/blog/incident-response-for-ai-systems/</guid><description>A response playbook for quality regressions, prompt attacks, retrieval contamination, runaway agents, cost spikes, and provider failures.</description><pubDate>Tue, 30 Jun 2026 00:00:00 GMT</pubDate><category>Product Architecture</category><category>incident response</category><category>observability</category><category>agents</category><category>LLMOps</category><author>Vishal Rajpurohit</author></item><item><title>Production RAG Is an Evidence Pipeline, Not a Vector Search</title><link>https://vishal.life/blog/production-rag-evidence-pipeline/</link><guid isPermaLink="true">https://vishal.life/blog/production-rag-evidence-pipeline/</guid><description>Designing retrieval-augmented generation around ingestion quality, query planning, evidence assembly, and verifiable answers.</description><pubDate>Mon, 29 Jun 2026 00:00:00 GMT</pubDate><category>Retrieval</category><category>RAG</category><category>retrieval</category><category>grounding</category><category>knowledge systems</category><author>Vishal Rajpurohit</author></item><item><title>Hybrid Search: Making Lexical and Vector Retrieval Cooperate</title><link>https://vishal.life/blog/hybrid-search-rank-fusion-engineering/</link><guid isPermaLink="true">https://vishal.life/blog/hybrid-search-rank-fusion-engineering/</guid><description>A practical design for candidate generation, score normalization, rank fusion, metadata filters, and retrieval evaluation.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate><category>Retrieval</category><category>hybrid search</category><category>BM25</category><category>vector search</category><category>ranking</category><author>Vishal Rajpurohit</author></item><item><title>Rerankers: The Quality Layer Between Search and Generation</title><link>https://vishal.life/blog/rerankers-the-quality-layer-after-retrieval/</link><guid isPermaLink="true">https://vishal.life/blog/rerankers-the-quality-layer-after-retrieval/</guid><description>How cross-encoders, late interaction, listwise ranking, and calibration turn noisy candidates into usable evidence.</description><pubDate>Fri, 12 Jun 2026 00:00:00 GMT</pubDate><category>Retrieval</category><category>reranking</category><category>cross-encoders</category><category>retrieval</category><category>relevance</category><author>Vishal Rajpurohit</author></item><item><title>Engineering the Agent Tool Loop</title><link>https://vishal.life/blog/agent-tool-loop-engineering/</link><guid isPermaLink="true">https://vishal.life/blog/agent-tool-loop-engineering/</guid><description>A concrete architecture for tool selection, state transitions, permissions, budgets, and recovery in production agents.</description><pubDate>Fri, 05 Jun 2026 00:00:00 GMT</pubDate><category>Agents</category><category>agents</category><category>tool use</category><category>orchestration</category><category>state machines</category><author>Vishal Rajpurohit</author></item><item><title>Durable Execution for Agents That Outlive a Web Request</title><link>https://vishal.life/blog/durable-execution-for-long-running-ai-agents/</link><guid isPermaLink="true">https://vishal.life/blog/durable-execution-for-long-running-ai-agents/</guid><description>Persisting checkpoints, replaying deterministically, handling human approval, and surviving crashes without duplicating side effects.</description><pubDate>Wed, 27 May 2026 00:00:00 GMT</pubDate><category>Agents</category><category>durable execution</category><category>workflows</category><category>agents</category><category>reliability</category><author>Vishal Rajpurohit</author></item><item><title>Build an LLM Evaluation Harness That Can Block a Release</title><link>https://vishal.life/blog/llm-evaluation-harness-that-guides-releases/</link><guid isPermaLink="true">https://vishal.life/blog/llm-evaluation-harness-that-guides-releases/</guid><description>From task contracts and test slices to calibrated judges, regression budgets, and production feedback loops.</description><pubDate>Mon, 18 May 2026 00:00:00 GMT</pubDate><category>Evaluation and Security</category><category>evals</category><category>LLM testing</category><category>quality</category><category>release engineering</category><author>Vishal Rajpurohit</author></item><item><title>Hallucination Controls Belong in the System, Not One Prompt</title><link>https://vishal.life/blog/hallucination-controls-built-into-the-system/</link><guid isPermaLink="true">https://vishal.life/blog/hallucination-controls-built-into-the-system/</guid><description>A layered design for constraining claims, grounding evidence, verifying outputs, and abstaining when the system does not know.</description><pubDate>Sat, 09 May 2026 00:00:00 GMT</pubDate><category>Evaluation and Security</category><category>hallucinations</category><category>grounding</category><category>verification</category><category>reliability</category><author>Vishal Rajpurohit</author></item><item><title>Guardrails as Policy Enforcement, Not Keyword Blocking</title><link>https://vishal.life/blog/guardrails-as-policy-enforcement/</link><guid isPermaLink="true">https://vishal.life/blog/guardrails-as-policy-enforcement/</guid><description>Building layered input, action, and output controls that remain testable when language and threats change.</description><pubDate>Thu, 30 Apr 2026 00:00:00 GMT</pubDate><category>Evaluation and Security</category><category>guardrails</category><category>policy</category><category>safety</category><category>moderation</category><author>Vishal Rajpurohit</author></item><item><title>Threat Modeling an AI Application End to End</title><link>https://vishal.life/blog/threat-modeling-ai-applications/</link><guid isPermaLink="true">https://vishal.life/blog/threat-modeling-ai-applications/</guid><description>Assets, trust boundaries, model-specific attacks, tool abuse, data leakage, and concrete mitigations for deployed AI systems.</description><pubDate>Tue, 21 Apr 2026 00:00:00 GMT</pubDate><category>Evaluation and Security</category><category>AI security</category><category>threat modeling</category><category>least privilege</category><category>data protection</category><author>Vishal Rajpurohit</author></item><item><title>Prompt Injection Defense in Depth</title><link>https://vishal.life/blog/prompt-injection-defense-in-depth/</link><guid isPermaLink="true">https://vishal.life/blog/prompt-injection-defense-in-depth/</guid><description>Why delimiters are not a sandbox, how indirect injection reaches agents, and which architectural controls actually reduce impact.</description><pubDate>Sun, 12 Apr 2026 00:00:00 GMT</pubDate><category>Evaluation and Security</category><category>prompt injection</category><category>agents</category><category>security</category><category>tool safety</category><author>Vishal Rajpurohit</author></item><item><title>Embedding Systems: Models, Indexes, and Semantic Drift</title><link>https://vishal.life/blog/embedding-systems-models-indexes-drift/</link><guid isPermaLink="true">https://vishal.life/blog/embedding-systems-models-indexes-drift/</guid><description>Selecting representations, constructing embedding inputs, migrating indexes, and detecting when semantic neighborhoods stop serving the product.</description><pubDate>Fri, 03 Apr 2026 00:00:00 GMT</pubDate><category>Retrieval</category><category>embeddings</category><category>vector search</category><category>semantic search</category><category>drift</category><author>Vishal Rajpurohit</author></item><item><title>Fine-Tuning with LoRA: Data, Adapters, and Production Operations</title><link>https://vishal.life/blog/fine-tuning-lora-and-adapter-operations/</link><guid isPermaLink="true">https://vishal.life/blog/fine-tuning-lora-and-adapter-operations/</guid><description>When fine-tuning is justified, how low-rank adapters work, and what it takes to evaluate, serve, and update them safely.</description><pubDate>Wed, 25 Mar 2026 00:00:00 GMT</pubDate><category>Inference Systems</category><category>fine-tuning</category><category>LoRA</category><category>adapters</category><category>training data</category><author>Vishal Rajpurohit</author></item><item><title>Multimodal AI Pipelines for Images, Audio, and Documents</title><link>https://vishal.life/blog/multimodal-ai-pipelines-images-audio-documents/</link><guid isPermaLink="true">https://vishal.life/blog/multimodal-ai-pipelines-images-audio-documents/</guid><description>Engineering preprocessing, temporal and spatial grounding, context budgets, validation, and storage around multimodal models.</description><pubDate>Mon, 16 Mar 2026 00:00:00 GMT</pubDate><category>Multimodal and Edge</category><category>multimodal</category><category>vision</category><category>audio</category><category>document AI</category><author>Vishal Rajpurohit</author></item><item><title>LLM Cost Engineering: From Token Prices to Unit Economics</title><link>https://vishal.life/blog/llm-cost-engineering-unit-economics/</link><guid isPermaLink="true">https://vishal.life/blog/llm-cost-engineering-unit-economics/</guid><description>Modeling full request cost, eliminating waste, forecasting margins, and optimizing without disguising quality regressions.</description><pubDate>Sat, 07 Mar 2026 00:00:00 GMT</pubDate><category>Inference Systems</category><category>LLM cost</category><category>unit economics</category><category>optimization</category><category>FinOps</category><author>Vishal Rajpurohit</author></item><item><title>Observability for LLM and Agent Systems</title><link>https://vishal.life/blog/observability-for-llm-and-agent-systems/</link><guid isPermaLink="true">https://vishal.life/blog/observability-for-llm-and-agent-systems/</guid><description>Tracing model calls, retrieval, tools, state transitions, quality signals, and cost without turning telemetry into a privacy liability.</description><pubDate>Thu, 26 Feb 2026 00:00:00 GMT</pubDate><category>Product Architecture</category><category>observability</category><category>tracing</category><category>agents</category><category>operations</category><author>Vishal Rajpurohit</author></item><item><title>Model Routing, Cascades, and Fallbacks</title><link>https://vishal.life/blog/model-routing-cascades-and-fallbacks/</link><guid isPermaLink="true">https://vishal.life/blog/model-routing-cascades-and-fallbacks/</guid><description>How to choose models per request using task risk, calibrated confidence, operational health, and total expected cost.</description><pubDate>Tue, 17 Feb 2026 00:00:00 GMT</pubDate><category>Product Architecture</category><category>model routing</category><category>cascades</category><category>fallbacks</category><category>inference</category><author>Vishal Rajpurohit</author></item><item><title>Structured Outputs Beyond Valid JSON</title><link>https://vishal.life/blog/structured-outputs-beyond-valid-json/</link><guid isPermaLink="true">https://vishal.life/blog/structured-outputs-beyond-valid-json/</guid><description>Designing schemas, constrained decoding, semantic validators, repairs, and safe evolution for dependable model integrations.</description><pubDate>Sun, 08 Feb 2026 00:00:00 GMT</pubDate><category>Product Architecture</category><category>structured outputs</category><category>JSON Schema</category><category>validation</category><category>API design</category><author>Vishal Rajpurohit</author></item><item><title>Model Context Protocol in Production</title><link>https://vishal.life/blog/model-context-protocol-production-architecture/</link><guid isPermaLink="true">https://vishal.life/blog/model-context-protocol-production-architecture/</guid><description>Designing MCP servers, capability discovery, authorization, transport, schemas, and observability for real agent ecosystems.</description><pubDate>Fri, 30 Jan 2026 00:00:00 GMT</pubDate><category>Agents</category><category>MCP</category><category>tools</category><category>protocols</category><category>agents</category><author>Vishal Rajpurohit</author></item><item><title>Edge LLM Inference: Designing for the Device</title><link>https://vishal.life/blog/edge-llm-inference-on-device/</link><guid isPermaLink="true">https://vishal.life/blog/edge-llm-inference-on-device/</guid><description>Quantization, memory, thermal budgets, runtimes, privacy, and hybrid execution for useful models on constrained hardware.</description><pubDate>Wed, 21 Jan 2026 00:00:00 GMT</pubDate><category>Multimodal and Edge</category><category>edge inference</category><category>quantization</category><category>on-device AI</category><category>privacy</category><author>Vishal Rajpurohit</author></item><item><title>AI Data Flywheels Without Feedback Pollution</title><link>https://vishal.life/blog/ai-data-flywheels-without-feedback-pollution/</link><guid isPermaLink="true">https://vishal.life/blog/ai-data-flywheels-without-feedback-pollution/</guid><description>Turning product use into better data through instrumentation, review, provenance, counterfactuals, and controlled training loops.</description><pubDate>Mon, 12 Jan 2026 00:00:00 GMT</pubDate><category>Retrieval</category><category>data flywheel</category><category>feedback</category><category>training data</category><category>ML operations</category><author>Vishal Rajpurohit</author></item><item><title>AI Product UX for Systems That Can Be Wrong</title><link>https://vishal.life/blog/ai-product-ux-for-uncertain-systems/</link><guid isPermaLink="true">https://vishal.life/blog/ai-product-ux-for-uncertain-systems/</guid><description>Designing intent, progress, evidence, editing, confirmation, and recovery so capability becomes trustworthy product behavior.</description><pubDate>Sun, 28 Dec 2025 00:00:00 GMT</pubDate><category>Product Architecture</category><category>AI UX</category><category>product design</category><category>trust</category><category>agents</category><author>Vishal Rajpurohit</author></item><item><title>Production AI Architecture in Laravel and PHP</title><link>https://vishal.life/blog/laravel-php-production-ai-architecture/</link><guid isPermaLink="true">https://vishal.life/blog/laravel-php-production-ai-architecture/</guid><description>Queues, streaming, typed provider boundaries, retrieval, tool execution, and observability for durable AI features in Laravel.</description><pubDate>Tue, 16 Dec 2025 00:00:00 GMT</pubDate><category>Product Architecture</category><category>Laravel</category><category>PHP</category><category>AI engineering</category><category>queues</category><author>Vishal Rajpurohit</author></item></channel></rss>