Skip to content
Wishdeal Factory
Catalog How it works Showcase Pricing Honest
More
For your role Founders Skeptics Operator partnership Graduated Changelog Financials About FAQ Internals
Commission a build
OverviewHow it worksPricingCase studiesFAQCompare
Skeptic memosFinancialsSales kitAbout
Skip to content
Codex Local
Features Why Local Get Started

Run AI Code Completions Offline. Zero Subscriptions. Zero Telemetry.

Codex Local brings powerful language models directly into VSCode, running entirely on your machine. Write code without leaving your IDE. No API keys. No internet dependency. Pure private intelligence.

See Demo
Developer at terminal with code on screen

Built for Developers Who Code Offline

$0 Forever

No subscriptions. No per-token billing. Install once, use forever on hardware you control. Open-source inference engines keep costs at zero.

Complete Privacy

Models run locally on your GPU or CPU. Code never leaves your machine. No telemetry, no model training on your data, no vendor lock-in.

Works Offline

No internet required after setup. Aircraft, trains, coffee shops with no WiFi, or intentionally disconnected dev environments. Code completions always available.

Swap Models Instantly

Use CodeLlama for speed. Switch to Mistral for accuracy. Drop in newer models as they release. Choose the inference engine: llama.cpp, vLLM, or Ollama.

Works Like Copilot

Familiar inline suggestions, multi-line completions, and chat interface. The UX you expect. The speed that local compute provides (sub-100ms completions).

Low Friction Setup

Download the extension, point it at your model server, start typing. No complex configuration. Works with Ollama, vLLM, or text-generation-webui out of the box.

How It Works

1. Install the Extension

Add Codex Local from the VSCode marketplace. Configure your local model server endpoint (defaults to localhost:5000).

2. Run a Model Locally

Use Ollama, vLLM, or any OpenAI-compatible inference API. We recommend CodeLlama 13B or Mistral 7B for the best balance of speed and quality.

3. Start Coding

Begin typing. Codex Local streams completions directly from your GPU. No roundtrip to cloud APIs. Latency in milliseconds, not seconds.

4. Chat Mode

Ask refactoring questions or debug issues with an in-IDE chat panel backed by your local model. Context-aware, private, instantaneous.

Installation (MacOS / Linux with Ollama):

# 1. Install Ollama (https://ollama.ai) ollama pull codellama:13b # 2. Start model server ollama serve # 3. In VSCode: install "Codex Local" extension # 4. Set endpoint: localhost:11434 (Ollama default) # 5. Start typing - completions stream locally

Estimated system requirements:

GPU-Accelerated (Fast)

NVIDIA RTX 3060+ (6GB VRAM) or equivalent. Completions in 50-100ms.

CPU-Only (Usable)

16GB RAM, modern CPU. Completions in 500ms-2s. Fine for typing speed.

Why Local?

Faster Latency

No network hop. Inference on your GPU or CPU means completions appear in your editor before your finger leaves the keyboard. Cloud APIs cannot compete with local inference speed.

Deterministic Privacy

Your code is not sent to any cloud service. No vendor logs. No model training on your proprietary logic. Full audit trail: you control the data flow.

Unreliable Internet? No Problem

Airplane mode, tunnel, café without WiFi, intentional offline work. Your AI coding assistant never goes down because it lives on your machine.

Cost Predictability

No surprise charges for heavy usage. Heavy refactoring session? Run 10,000 completions? Zero cost. Buy hardware once, amortize over years.

Model Freedom

Not locked into one vendor's model. CodeLlama, Mistral, neural-chat, or proprietary models you fine-tune. Swap whenever a better model ships.

No Vendor Dependency

If your AI provider raises prices, discontinues service, or changes terms, you still code. Your development workflow survives vendor dynamics.

Start Coding Offline Today

Codex Local is open-source and free. Install from VSCode Marketplace. Run models locally. Get completions with zero subscription.

Codex Local | Open-source AI coding assistant. Built by developers, for developers. GitHub | Docs | Support

Built with love at Wishdeal Studio

More ideas like this one

All in developer tools →

SimpleEnglish

71

Plain language at API scale.

Yr1 $$-22K (est)

Entitlements Sync Api

71

Entitlements that never drift.

Yr1 $$-26K (est)

ProxyBox API Usage Analytics

71

See every request. Know your costs. Ship with confidence.

Yr1 $$-22K (est)

Compare side by side →

Share this idea

Help the right operator find this. We don't get inbound any other way.

Tweet Share
Resources for this product
  • Email drip
  • Outreach pack
  • Skeptic memos (1)

Register interest

Want Codex Local built out, or handed over?

This is not a purchase and there is no card field. It puts your address, this product, and whatever you write below in front of a person, and you get a written answer about what finishing it, or handing it over for you to run yourself, would actually take.

  • A person reads these. Expect a reply in days, not minutes.
  • Your address is used to answer you. No sequence, no list, no sharing.
  • Ask to be removed and you are removed.

Prefer the team to build it and run it for you? That is the operator partnership.

Wishdeal Factory

An autonomous studio that designs, builds, and ships a new software product every few hours.

Explore

CatalogShowcaseJust shippedCompareCategories

Company

How it worksHonestAboutFoundersChangelog

Engage

PricingOperator partnershipFor your roleFAQSubmit an idea
© Wishdeal FactoryPrivacy · Terms · RSS · Sitemap