Tagged "rag"
- Liquid AI Releases LFM2.5-VL-3B: Compact Vision-Language Model for Edge Inference
- Liquid AI LFM2.5-2.6B: Open-Weights Agentic Model With 128K Context and Tool Calling
- K-EXAONE 2.0 Brings 262K Context to Frontier AI
- Stop Paying for Search APIs—This Self-Hosted Tool Lets Your Local LLM Search the Web for Free
- Onemind.md – Adding Repository Memory to LLMs Without Extra Tooling
- SigMap: 97% Token Reduction for AI Coding Sessions
- code-on-incus: Isolated Machine Environments for AI Agents
- Building a Personal Ebook Librarian with Local LLMs for Better Recommendations
- LongCat-2.0 Released
- Beyond Setup: Production Practices for Local LLM Deployment
- Meet EverOS: An Open Source Markdown-First Agent Memory Runtime With Hybrid BM25 + Vector Retrieval
- LLM-Free, Layout-Aware PDF Chunker in Pure Rust
- Local Semantic Search Engine in Rust, No External DB
- You Can Now Run Max AI Models on Apple Silicon
- GEEKOM A9 Max Delivers 32GB RAM and Native LLM Support in Compact Form Factor
- Qwable: New Free Local Model Brings Claude-like Capabilities to Edge Devices
- Mac Mini Positioned as Premier On-Device AI Computer for Local LLM Inference
- PageToMD – A CLI tool to turn web pages into clean Markdown for AI agents
- Architecting Modular Local AI Ecosystems to Escape Token Economics
- Show HN: Veritrooper – find what your AI gets wrong about your own docs
- Ask HN: What is the AI setup for an experienced dev starting on a new project?
- Show HN: LLM Memory Without Context Bleed – 100% Precision vs. <10% Vector Search
- Run Llama.cpp In-Process from Java with Project Panama FFM
- LLM Memory Systems Benchmark: High Recall, Near-Zero Precision for Tested Systems
- Local LLM Setup: How to Use RAG and an Embedding Model to Stop Wasting Context
- LLM Wiki App Chunker: Transform Documents Into Navigable Knowledge Trees
- On-Device AI to Be in 80% of Wearables by 2032
- Local LLM Integration Enables Replacement of Paid Subscription Services
- Discussion: Including New Mathematical Proofs in LLM Training Data for Rediscovery
- Agentic AI Community Focus: Building Local Agents in 2026
- Show HN: Memex, Claude Memory via Local RAG with MCP and Offline Embeddings
- SQL Server 2025 Adds Built-in Chunking and Vector Support
- Xmemory: Benchmarking Structured AI Memory Against RAG and Hybrid RAG
- N8n, Dify, and Ollama Might Be the Best Self-Hosted AI Automation Stack Right Now
- Mathesar 0.10.0
- 16 Ways to Make a Small Language Model Think Bigger
- N8n, Dify, and Ollama Emerge as Leading Self-Hosted AI Automation Stack
- Minisforum N5 MAX AI NAS Delivers 126 TOPS with 200TB Storage for Local LLM Workloads
- Context Window Optimization: Extending Gemma 4 Context Length Through Efficient Projection Quantization
- Does RAG Help AI Coding Tools?
- Ask HN: What do you use for local embeddings?
- RAG Deployment Lessons from Regulated Industries
- Lat.md: Agent Lattice – A Knowledge Graph for Your Codebase in Markdown
- LM Studio Releases Reworked Plugins with Fully Local Web Research
- Velr: Embedded Property-Graph Database for Local LLM Applications
- Powerful AI Search Engine Built on Single GeForce RTX 5090
- Llama 8B Matches 70B Performance on Multi-Hop QA Using Structured Prompting
- LMCache Dramatically Accelerates LLM Inference on Oracle Data Science Platform
- MiniMax-M2.7: New Compact Model Announced for Local Deployment
- Mamba 3: State Space Model Architecture Optimized for Inference
- Gloss: Open-Source, Local-First RAG Alternative to NotebookLM Built in Rust
- AI Agent Reliability Tracker
- Framework Choice Critical: llama.cpp and vLLM Outperform Ollama for Qwen 3.5 Testing
- RAG vs. Skill vs. MCP vs. RLM: Comparing LLM Enhancement Patterns
- RAG-Enterprise – 100% Local RAG System for Enterprise Documents
- Building a Privacy-Preserving RAG System in the Browser
- Researchers Develop Persistent Memory System for Local LLMs—No RAG Required
- Elastic Introduces Best-in-Class Embedding Models for High Performance Semantic Search
- Elastic Introduces Best-in-Class Embedding Models for High Performance Semantic Search
- Search and Analyze Documents from the DOJ Epstein Files Release with Local LLM
- NVIDIA Releases Dynamo v0.9.0: Infrastructure Overhaul With FlashIndexer and Multi-Modal Support
- Local-First RAG: Vector Search in SQLite with Hamming Distance
- InitRunner: YAML-Based AI Agent Framework with RAG and Memory
- GPU-Accelerated DataFrame Library for Local Inference Workloads
- Microsoft MarkItDown: Document Preprocessing Tool for LLMs
- Building a RAG Pipeline on 2M+ Pages: EpsteinFiles-RAG Project