选择你的计算附加组件
Choosing the right Compute Add-on for your vector workload.
你有两种方式来扩展你的向量工作负载:
🌐 You have two options for scaling your vector workload:
- 增加你的数据库大小。这份指南会帮你为你的工作负载选择合适的大小。
- 将你的工作负载分散到多个数据库上。你可以在《Engineering for Scale》(engineering-for-scale)中找到关于这种方法的更多细节。
维度 #
🌐 Dimensionality
在选择合适的计算插件时,你的嵌入向量维度是最重要的因素。一般来说,维度越低,性能越好。下面我们提供了一些常见嵌入维度的指南。对于每个基准测试,我们使用 Vecs 创建了一个集合,将嵌入上传到单个表中,并为嵌入列创建了 inner-product 距离度量的 IVFFlat 和 HNSW 索引。然后我们运行了一系列查询来衡量不同计算插件的性能:
🌐 The number of dimensions in your embeddings is the most important factor in choosing the right Compute Add-on. In general, the lower the dimensionality the better the performance. We've provided guidance for some of the more common embedding dimensions below. For each benchmark, we used Vecs to create a collection, upload the embeddings to a single table, and create both the IVFFlat and HNSW indexes for inner-product distance measure for the embedding column. We then ran a series of queries to measure the performance of different compute add-ons:
HNSW#
384 维 #
🌐 384 dimensions [#hnsw-384-dimensions]
这个基准测试使用了 dbpedia-entities-openai-1M 数据集,其中包含 1,000,000 个文本向量,已重新生成为 384 维度的向量。每个向量都是使用 gte-small 生成的。
🌐 This benchmark uses the dbpedia-entities-openai-1M dataset containing 1,000,000 embeddings of text, regenerated for 384 dimension embeddings. Each embedding is generated using gte-small.
| 计算规模 | 向量数 | m | 构建ef | 搜索ef | 查询每秒 (QPS) | 平均延迟 | 95%延迟 | RAM 使用 | RAM |
|---|---|---|---|---|---|---|---|---|---|
| 微型 | 100,000 | 16 | 64 | 60 | 580 | 0.017 秒 | 0.024 秒 | 1.2 (交换) | 1 GB |
| 小型 | 250,000 | 24 | 64 | 60 | 440 | 0.022 秒 | 0.033 秒 | 2 GB | 2 GB |
| 中型 | 500,000 | 24 | 64 | 80 | 350 | 0.028 秒 | 0.045 秒 | 4 GB | 4 GB |
| 大型 | 1,000,000 | 32 | 80 | 100 | 270 | 0.073 秒 | 0.108 秒 | 7 GB | 8 GB |
| XL | 1,000,000 | 32 | 80 | 100 | 525 | 0.038 秒 | 0.059 秒 | 9 GB | 16 GB |
| 2XL | 1,000,000 | 32 | 80 | 100 | 790 | 0.025 秒 | 0.037 秒 | 9 GB | 32 GB |
| 4XL | 1,000,000 | 32 | 80 | 100 | 1650 | 0.015 秒 | 0.018 秒 | 11 GB | 64 GB |
| 8XL | 1,000,000 | 32 | 80 | 100 | 2690 | 0.015 秒 | 0.016 秒 | 13 GB | 128 GB |
| 12XL | 1,000,000 | 32 | 80 | 100 | 3900 | 0.014 秒 | 0.016 秒 | 13 GB | 192 GB |
| 16XL | 1,000,000 | 32 | 80 | 100 | 4200 | 0.014 秒 | 0.016 秒 | 20 GB | 256 GB |
基准测试的准确率是0.99。
🌐 Accuracy was 0.99 for benchmarks.
960 维 #
🌐 960 dimensions [#hnsw-960-dimensions]
这个基准使用了 gist-960 数据集,其中包含 1,000,000 个图片的嵌入。每个嵌入有 960 个维度。
🌐 This benchmark uses the gist-960 dataset, which contains 1,000,000 embeddings of images. Each embedding is 960 dimensions.
| 计算大小 | 向量 | m | 构建效率 ef_construction | 搜索效率 ef_search | 每秒查询数 QPS | 平均延迟 | 95百分位延迟 | 内存使用 | 内存 |
|---|---|---|---|---|---|---|---|---|---|
| 微型 | 30,000 | 16 | 64 | 65 | 430 | 0.024 秒 | 0.034 秒 | 1.2 GB(交换) | 1 GB |
| 小 | 100,000 | 32 | 80 | 60 | 260 | 0.040 秒 | 0.054 秒 | 2.2 GB(交换) | 2 GB |
| 中等 | 250,000 | 32 | 80 | 90 | 120 | 0.083 秒 | 0.106 秒 | 4 GB | 4 GB |
| 大 | 500,000 | 32 | 80 | 120 | 160 | 0.063 秒 | 0.087 秒 | 7 GB | 8 GB |
| XL | 1,000,000 | 32 | 80 | 200 | 200 | 0.049 秒 | 0.072 秒 | 13 GB | 16 GB |
| 2XL | 1,000,000 | 32 | 80 | 200 | 340 | 0.025 秒 | 0.029 秒 | 17 GB | 32 GB |
| 4XL | 1,000,000 | 32 | 80 | 200 | 630 | 0.031 秒 | 0.050 秒 | 18 GB | 64 GB |
| 8XL | 1,000,000 | 32 | 80 | 200 | 1100 | 0.034 秒 | 0.048 秒 | 19 GB | 128 GB |
| 12XL | 1,000,000 | 32 | 80 | 200 | 1420 | 0.041 秒 | 0.095 秒 | 21 GB | 192 GB |
| 16XL | 1,000,000 | 32 | 80 | 200 | 1650 | 0.037 秒 | 0.081 秒 | 23 GB | 256 GB |
基准测试的准确率是0.99。
🌐 Accuracy was 0.99 for benchmarks.
通过增加m和ef_construction也可以提高 QPS。这将允许你使用更小的ef_search值,从而提高 QPS。
🌐 QPS can also be improved by increasing m and ef_construction. This will allow you to use a smaller value for ef_search and increase QPS.
1536 维度 #
🌐 1536 dimensions [#hnsw-1536-dimensions]
这个基准使用了 dbpedia-entities-openai-1M 数据集,其中包含 1,000,000 个文本嵌入。同时,对于计算附加组件 large 及以下,还使用了 224,482 个来自 维基百科文章 的嵌入。每个嵌入都是 1536 维,由 OpenAI Embeddings API 创建。
🌐 This benchmark uses the dbpedia-entities-openai-1M dataset, which contains 1,000,000 embeddings of text. And 224,482 embeddings from Wikipedia articles for compute add-ons large and below. Each embedding is 1536 dimensions created with the OpenAI Embeddings API.
| 计算大小 | 向量 | m | 构建效率 ef_construction | 搜索效率 ef_search | 每秒查询数 QPS | 平均延迟 | 95百分位延迟 | 内存使用 | 内存 |
|---|---|---|---|---|---|---|---|---|---|
| 微型 | 15,000 | 16 | 40 | 40 | 480 | 0.011 秒 | 0.016 秒 | 1.2 GB(交换) | 1 GB |
| 小 | 50,000 | 32 | 64 | 100 | 175 | 0.031 秒 | 0.051 秒 | 2.2 GB(交换) | 2 GB |
| 中等 | 100,000 | 32 | 64 | 100 | 240 | 0.083 秒 | 0.126 秒 | 4 GB | 4 GB |
| 大 | 224,482 | 32 | 64 | 100 | 280 | 0.017 秒 | 0.028 秒 | 8 GB | 8 GB |
| XL | 500,000 | 24 | 56 | 100 | 360 | 0.055 秒 | 0.135 秒 | 13 GB | 16 GB |
| 2XL | 1,000,000 | 24 | 56 | 250 | 560 | 0.036 秒 | 0.058 秒 | 32 GB | 32 GB |
| 4XL | 1,000,000 | 24 | 56 | 250 | 950 | 0.021 秒 | 0.033 秒 | 39 GB | 64 GB |
| 8XL | 1,000,000 | 24 | 56 | 250 | 1650 | 0.016 秒 | 0.023 秒 | 40 GB | 128 GB |
| 12XL | 1,000,000 | 24 | 56 | 250 | 1900 | 0.015 秒 | 0.021 秒 | 38 GB | 192 GB |
| 16XL | 1,000,000 | 24 | 56 | 250 | 2200 | 0.015 秒 | 0.020 秒 | 40 GB | 256 GB |
基准测试的准确率是0.99。
🌐 Accuracy was 0.99 for benchmarks.
QPS 也可以通过增加 m 和 ef_construction 来提升。这将允许你使用更小的 ef_search 值并增加 QPS。例如,将 4XL 的 m 提高到 32,ef_construction 提高到 80,QPS 将增加到 1280。
🌐 QPS can also be improved by increasing m and ef_construction. This will allow you to use a smaller value for ef_search and increase QPS. For example, increasing m to 32 and ef_construction to 80 for 4XL will increase QPS to 1280.
如果内存允许,可以向同一个表上传更多向量(例如,OpenAI 嵌入的 4XL 及以上计划)。但这会影响查询性能:QPS 会下降,延迟会增加。扩展应该几乎是线性的,但建议对你的工作负载进行基准测试,以找到每个表和每个数据库实例的最优向量数量。
🌐 It is possible to upload more vectors to a single table if Memory allows it (for example, 4XL plan and higher for OpenAI embeddings). But it will affect the performance of the queries: QPS will be lower, and latency will be higher. Scaling should be almost linear, but it is recommended to benchmark your workload to find the optimal number of vectors per table and per database instance.
下面的图表比较了在不同嵌入维度下,各种计算规模的 HNSW 每秒查询次数。
🌐 The chart below compares HNSW queries-per-second across compute sizes for different embedding dimensions.

体外受精平板 #
🌐 IVFFlat
384 维 #
🌐 384 dimensions [#ivfflat-384-dimensions]
这个基准测试使用了 dbpedia-entities-openai-1M 数据集,其中包含 1,000,000 个文本向量,已重新生成为 384 维度的向量。每个向量都是使用 gte-small 生成的。
🌐 This benchmark uses the dbpedia-entities-openai-1M dataset containing 1,000,000 embeddings of text, regenerated for 384 dimension embeddings. Each embedding is generated using gte-small.
| 计算规模 | 向量 | 列表 | 探针 | 每秒查询数 (QPS) | 平均延迟 | 95百分位延迟 | 内存使用 | 内存 |
|---|---|---|---|---|---|---|---|---|
| 微型 | 100,000 | 500 | 50 | 205 | 0.048 秒 | 0.066 秒 | 1.2 GB (交换) | 1 GB |
| 小型 | 250,000 | 1000 | 60 | 160 | 0.062 秒 | 0.079 秒 | 2 GB | 2 GB |
| 中型 | 500,000 | 2000 | 80 | 120 | 0.082 秒 | 0.104 秒 | 3.2 GB | 4 GB |
| 大型 | 1,000,000 | 5000 | 150 | 75 | 0.269 秒 | 0.375 秒 | 6.5 GB | 8 GB |
| 超大 | 1,000,000 | 5000 | 150 | 150 | 0.131 秒 | 0.178 秒 | 9 GB | 16 GB |
| 2倍超大 | 1,000,000 | 5000 | 150 | 300 | 0.066 秒 | 0.099 秒 | 10 GB | 32 GB |
| 4倍超大 | 1,000,000 | 5000 | 150 | 570 | 0.035 秒 | 0.046 秒 | 10 GB | 64 GB |
| 8倍超大 | 1,000,000 | 5000 | 150 | 1400 | 0.023 秒 | 0.028 秒 | 12 GB | 128 GB |
| 12倍超大 | 1,000,000 | 5000 | 150 | 1550 | 0.030 秒 | 0.039 秒 | 12 GB | 192 GB |
| 16倍超大 | 1,000,000 | 5000 | 150 | 1800 | 0.030 秒 | 0.039 秒 | 16 GB | 256 GB |
960 维 #
🌐 960 dimensions [#ivfflat-960-dimensions]
这个基准使用了 gist-960 数据集,其中包含 1,000,000 个图片的嵌入。每个嵌入有 960 个维度。
🌐 This benchmark uses the gist-960 dataset, which contains 1,000,000 embeddings of images. Each embedding is 960 dimensions.
| 计算规模 | 向量 | 列表 | QPS | 平均延迟 | 延迟 p95 | 内存使用 | 内存 |
|---|---|---|---|---|---|---|---|
| 微型 | 30,000 | 30 | 75 | 0.065 秒 | 0.088 秒 | 1.1 GB(交换) | 1 GB |
| 小型 | 100,000 | 100 | 78 | 0.064 秒 | 0.092 秒 | 1.8 GB | 2 GB |
| 中型 | 250,000 | 250 | 58 | 0.085 秒 | 0.129 秒 | 3.2 GB | 4 GB |
| 大型 | 500,000 | 500 | 55 | 0.088 秒 | 0.140 秒 | 5 GB | 8 GB |
| XL | 1,000,000 | 1000 | 110 | 0.046 秒 | 0.070 秒 | 14 GB | 16 GB |
| 2XL | 1,000,000 | 1000 | 235 | 0.083 秒 | 0.136 秒 | 10 GB | 32 GB |
| 4XL | 1,000,000 | 1000 | 420 | 0.071 秒 | 0.106 秒 | 11 GB | 64 GB |
| 8XL | 1,000,000 | 1000 | 815 | 0.072 秒 | 0.106 秒 | 13 GB | 128 GB |
| 12XL | 1,000,000 | 1000 | 1150 | 0.052 秒 | 0.078 秒 | 15.5 GB | 192 GB |
| 16XL | 1,000,000 | 1000 | 1345 | 0.072 秒 | 0.106 秒 | 17.5 GB | 256 GB |
1536 维度 #
🌐 1536 dimensions [#ivfflat-1536-dimensions]
这个基准测试使用了 dbpedia-entities-openai-1M 数据集,其中包含 1,000,000 个文本向量。每个向量有 1536 个维度,是通过 OpenAI Embeddings API 创建的。
🌐 This benchmark uses the dbpedia-entities-openai-1M dataset, which contains 1,000,000 embeddings of text. Each embedding is 1536 dimensions created with the OpenAI Embeddings API.
| 计算规模 | 向量数 | 列表数 | 每秒查询数 (QPS) | 平均延迟 | 95% 延迟 | 内存使用 | 内存总量 |
|---|---|---|---|---|---|---|---|
| 微型 | 20,000 | 40 | 135 | 0.372 秒 | 0.412 秒 | 1.2 GB(交换) | 1 GB |
| 小型 | 50,000 | 100 | 140 | 0.357 秒 | 0.398 秒 | 1.8 GB | 2 GB |
| 中型 | 100,000 | 200 | 130 | 0.383 秒 | 0.446 秒 | 3.7 GB | 4 GB |
| 大型 | 250,000 | 500 | 130 | 0.378 秒 | 0.434 秒 | 7 GB | 8 GB |
| 超大 (XL) | 500,000 | 1000 | 235 | 0.213 秒 | 0.271 秒 | 13.5 GB | 16 GB |
| 双超大 (2XL) | 1,000,000 | 2000 | 380 | 0.133 秒 | 0.236 秒 | 30 GB | 32 GB |
| 四超大 (4XL) | 1,000,000 | 2000 | 720 | 0.068 秒 | 0.120 秒 | 35 GB | 64 GB |
| 八超大 (8XL) | 1,000,000 | 2000 | 1250 | 0.039 秒 | 0.066 秒 | 38 GB | 128 GB |
| 十二超大 (12XL) | 1,000,000 | 2000 | 1600 | 0.030 秒 | 0.052 秒 | 41 GB | 192 GB |
| 十六超大 (16XL) | 1,000,000 | 2000 | 1790 | 0.029 秒 | 0.051 秒 | 45 GB | 256 GB |
对于 1,000,000 个向量,使用 10 个探针的准确率为 0.91。而对于 500,000 个向量及以下,使用 10 个探针的准确率在 0.95 到 0.99 之间。想要提高准确率,你需要增加探针的数量。
🌐 For 1,000,000 vectors 10 probes results to accuracy of 0.91. And for 500,000 vectors and below 10 probes results to accuracy in the range of 0.95 - 0.99. To increase accuracy, you need to increase the number of probes.
下面的图表绘制了每秒请求数与计算规模的关系。
🌐 The chart below plots requests-per-second against compute size.

如果内存允许,可以向同一个表上传更多向量(例如,OpenAI 嵌入的 4XL 及以上计划)。但这会影响查询性能:QPS 会下降,延迟会增加。扩展应该几乎是线性的,但建议对你的工作负载进行基准测试,以找到每个表和每个数据库实例的最优向量数量。
🌐 It is possible to upload more vectors to a single table if Memory allows it (for example, 4XL plan and higher for OpenAI embeddings). But it will affect the performance of the queries: QPS will be lower, and latency will be higher. Scaling should be almost linear, but it is recommended to benchmark your workload to find the optimal number of vectors per table and per database instance.
性能小贴士 #
🌐 Performance tips
有很多方法可以提升你的 pgvector 性能。这里有一些小贴士:
🌐 There are various ways to improve your pgvector performance. Here are some tips:
预热你的数据库 #
🌐 Pre-warming your database
在投入生产之前,先执行几千个“热身”查询是很有用的。这有助于 RAM 的使用。这也可以帮助你确定为你的工作负载选择的计算规模是否合适。
🌐 It's useful to execute a few thousand “warm-up” queries before going into production. This helps help with RAM utilization. This can also help to determine that you've selected the right compute size for your workload.
微调索引参数 #
🌐 Fine-tune index parameters
你可以通过增加 m、ef_construction 或 lists 来提高每秒请求数。不过这里有个重要提醒:这些参数的值越高,构建索引所需的时间也越长。
🌐 You can increase the Requests per Second by increasing m and ef_construction or lists. This also has an important caveat: building the index takes longer with higher values for these parameters.
下面的图表显示了 HNSW 构建参数 m 和 ef_construction 如何影响 dbpedia 数据集上的每秒请求数。
🌐 The chart below shows how the HNSW build parameters m and ef_construction affect requests-per-second on the dbpedia dataset.

在 Going to Production for AI applications 查看更多技巧和完整的逐步指南。
🌐 Check out more tips and the complete step-by-step guide in Going to Production for AI applications.
基准方法 #
🌐 Benchmark methodology
我们遵循ANN Benchmarks方法中概述的技术。一个 Python 测试运行器负责上传数据、创建索引以及执行查询。pgvector 引擎是使用 vecs 实现的,这是一个用于 pgvector 的 Python 客户端。
🌐 We follow techniques outlined in the ANN Benchmarks methodology. A Python test runner is responsible for uploading the data, creating the index, and running the queries. The pgvector engine is implemented using vecs, a Python client for pgvector.

上图显示了 vecs 基准测试的设置:一个 Python 测试运行器上传数据、构建索引,并对 pgvector 运行查询。
🌐 The diagram above shows the vecs benchmark setup: a Python test runner uploads data, builds the index, and runs queries against pgvector.
每次测试至少运行30-40分钟。测试包括在不同并发级别下进行的一系列实验,以测量引擎在不同负载类型下的性能。然后将结果取平均值。
🌐 Each test is run for a minimum of 30-40 minutes. They include a series of experiments executed at different concurrency levels to measure the engine's performance under different load types. The results are then averaged.
一般建议是,对于大多数工作负载,我们建议使用至少5的并发级别;对于高负载工作负载,则建议使用至少30的并发级别。
🌐 As a general recommendation, we suggest using a concurrency level of 5 or more for most workloads and 30 or more for high-load workloads.