Skip to content

Build1 publisher2 min readPublished

Full-precision embeddings push a 100-million-vector OpenSearch index to 1.3 TB of RAM

Amazon OpenSearch Service needs about 1.3 TB of resident RAM for 100 million 1,536-dimension FP32 vectors with one replica, a dev.to sizing post calculates. Raw vector values are about 98% of each entry, so the encoding picked before ingestion decides most of that memory and the node count behind it.

The Engineer · Build desk

Illustration accompanying Full-precision embeddings push a 100-million-vector OpenSearch index to 1.3 TB of RAM

What happened

  • Every footprint figure in the post is per copy, and one replica doubles it because a replica keeps a full second graph in memory.
  • The post prices the fleet needed to hold that index at roughly $31,000 to $46,000 a month on demand, before reserved-instance or savings-plan discounts.
  • Memory-optimized loading lets the index run in less RAM by paging vectors from disk on demand, at some cost in query latency.
  • Later parts of the series cover engine-side quantization and an on-disk mode that keeps a compressed copy in memory and rescans full-precision vectors from disk.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost At about $575 per on-demand instance-month, every 16 GB of vectors taken out of RAM on a default-configured cluster removes roughly one r8gd.2xlarge from the bill.
  • constraint Clusters that serve heavy text queries from the same nodes fall outside the 75% breaker advice, so their memory savings have to come from encoding or disk-backed loading.
  • exposure Client-side encoding puts the recall risk on the team, since OpenSearch Service searches whatever reduced-precision vectors arrive exactly as sent.

At the defaults, about a quarter of each node's RAM holds vectors. An r8gd.2xlarge has 64 GB. OpenSearch Service gives half to the JVM heap, and knn.memory.circuit_breaker.limit caps native vector memory at 50% of the 32 GB left off-heap, about 16 GB per instance [5][4]. Spread 1.3 TB across nodes at that rate and the count is roughly 80 [6].

The first cut the post offers is a setting. On a node dedicated to vector search, with no heavy text queries leaning on the heap, it raises the breaker to 75% of off-heap memory [7]. Usable vector memory rises to about 24 GB and the count falls to roughly 54 [7], about a third fewer nodes [2]. The vectors are stored exactly as before [5][7].

Precision acts on almost everything else. Each FP32 vector at 1,536 dimensions is 6,144 bytes, four per dimension [4]. The HNSW graph adds 8 bytes per neighbor link, 128 bytes at the default m of 16, and a 1.1 factor covers the index running about ten percent larger than its input [3]. Before that factor, the raw values are about 98% of each entry and the graph about 2% [1]. Tuning m works on the 2%. Encoding the values in fewer bytes works on the 98%, and the post calls precision the one choice, made before indexing, that sets how much memory you need [2].

The post's units are binary. At 6,899 bytes each, 100 million vectors come to 689.9 billion bytes [3][5]. In binary units that is 642.5 GiB, the post's 643 GB [5]. A capacity sheet that treats the figure as decimal gigabytes comes up about 7% short [6].

The post calls its dollar figures order-of-magnitude [8]. Both ends of its range divide out to about $575 per on-demand instance-month [3]. They carry over to another index only if the inputs match: 1,536 dimensions, m of 16, one replica, the default heap split and on-demand r8gd.2xlarge pricing [1][3][5][8].

In my view the post's order is right for a dedicated vector cluster. Take the breaker change first, because it leaves the vectors alone [7]. Then pick FP16, INT8 or binary encoding and measure what it does to recall, since the post frames each step down in precision as a measured trade of accuracy for memory [13].

What to watch

  • The series' installment on OpenSearch-native quantization, and whether engine-side compression matches client-side INT8 or binary encoding on recall at 100 million vectors.
  • The on-disk mode installment, and how much latency the full-precision second pass adds compared with a fully resident FP32 index.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories