Self-Hosted AI/LLM & RAG Platform Designed and built a self-hosted AI/LLM platform for inference,...Self-Hosted AI/LLM & RAG Platform Designed and built a self-hosted AI/LLM platform for inference,...
The network for creativity
Join 1.25M professional creatives like you
Connect with clients, get discovered, and run your business 100% commission-free
Creatives on Contra have earned over $150M and we are just getting started
Self-Hosted AI/LLM & RAG Platform Designed and built a self-hosted AI/LLM platform for inference, embeddings, and RAG workloads, with a focus on practical deployment, performance, and integration with production backend systems.
Deployed LLM inference using vLLM on NVIDIA GPUs and Ollama, configured GPU-enabled Docker environments, and evaluated CPU/GPU inference performance. Built RAG pipelines using embedding models, PostgreSQL vector search, document processing, and retrieval to provide relevant context to LLMs.
Integrated both self-hosted models and managed AI services through AWS Bedrock, exposing AI capabilities to existing backend applications and services.
Key technologies: vLLM, NVIDIA GPUs, Docker, Ollama, AWS Bedrock, PostgreSQL, vector search, embeddings, RAG, Python, LLM inference.
Post image
Back to feed
The network for creativity
Join 1.25M professional creatives like you
Connect with clients, get discovered, and run your business 100% commission-free
Creatives on Contra have earned over $150M and we are just getting started