Semantic Chunking is a strategy for splitting documents into smaller pieces (chunks) based on meaning, rather than on fixed character counts or simple separators. It tries to keep sentences or paragraphs that are about the same topic together, and split where the topic changes. Think of it as “a smart editor who knows where one idea ends and the next begins.
语义分块(Semantic Chunking)是一种基于语义(意思)将文档切分成小片段(chunk)的策略,而不是根据固定字符数或简单分隔符来切。它尽量把讨论同一个主题的句子或段落放在一起,在话题发生转折的地方切分。可以把它比喻为:“一个聪明的编辑,知道一个想法在哪里结束,下一个想法从哪里开始
We use an embedding model to measure the semantic similarity between consecutive sentences or small text segments. If the similarity drops below a threshold, we split at that point, creating a new chunk.
Code Example
import numpy as np
from typing import List, Tuple
import requests # 用于直接调用 embedding API,你也可以换成 openai 库
# ============================================================================
# 0. 配置部分 - 你可以换成自己的 API Key 和 Endpoint
# ============================================================================
API_KEY = "your-api-key"
ENDPOINT = "https://api.openai.com/v1/embeddings" # 或 Azure/DeepSeek 的 embedding endpoint
MODEL_NAME = "text-embedding-ada-002" # 或 text-embedding-3-small
# ============================================================================
# 1. 工具函数:获取单个文本的 embedding
# ============================================================================
def get_embedding(text: str) -> List[float]:
"""
Call the embedding API and return the embedding vector.
调用 Embedding API 并返回 embedding 向量。
"""
headers = {
"Authorization": f"Bearer {API_KEY}",
"Content-Type": "application/json",
}
payload = {
"input": text,
"model": MODEL_NAME,
}
resp = requests.post(ENDPOINT, headers=headers, json=payload)
resp.raise_for_status()
data = resp.json()
# 从返回的 JSON 中提取 embedding 向量
embedding = data["data"][0]["embedding"]
return embedding
# ============================================================================
# 2. 批量获取 embeddings(一次请求处理多个句子,节省 API 调用次数)
# ============================================================================
def get_embeddings_batch(texts: List[str]) -> List[List[float]]:
"""
Get embeddings for multiple texts in one API call.
一次性获取多个文本的 embedding。
"""
headers = {
"Authorization": f"Bearer {API_KEY}",
"Content-Type": "application/json",
}
payload = {
"input": texts,
"model": MODEL_NAME,
}
resp = requests.post(ENDPOINT, headers=headers, json=payload)
resp.raise_for_status()
data = resp.json()
# 按输入顺序提取所有 embedding
embeddings = [item["embedding"] for item in data["data"]]
return embeddings
# ============================================================================
# 3. 计算余弦相似度
# ============================================================================
def cosine_similarity(vec_a: List[float], vec_b: List[float]) -> float:
"""
Compute cosine similarity between two vectors.
计算两个向量的余弦相似度。
"""
a = np.array(vec_a)
b = np.array(vec_b)
dot_product = np.dot(a, b) # 点积
norm_a = np.linalg.norm(a) # L2 范数
norm_b = np.linalg.norm(b)
if norm_a == 0 or norm_b == 0:
return 0.0
return dot_product / (norm_a * norm_b) # 余弦相似度
# ============================================================================
# 4. 语义分块核心函数
# ============================================================================
def semantic_chunk(
document: str,
similarity_threshold: float = 0.8,
min_chunk_sentences: int = 3
) -> List[str]:
"""
Split document into semantic chunks based on embedding similarity.
根据 embedding 相似度将文档切分成语义块。
Args:
document: 输入文档字符串
similarity_threshold: 相似度阈值,低于此值则切分
min_chunk_sentences: 每个 chunk 至少包含的句子数
Returns:
分块后的字符串列表
"""
# ---- 4.1 简单分句(生产环境建议用 nltk/spaCy)----
# 以句号、问号、感叹号等进行分割,保留分隔符后处理
import re
raw_sentences = re.split(r'(?<=[.!?])\s+', document)
# 过滤掉空字符串
sentences = [s.strip() for s in raw_sentences if s.strip()]
if len(sentences) == 0:
return []
# ---- 4.2 获取所有句子的 embedding (批量)----
embeddings = get_embeddings_batch(sentences)
# ---- 4.3 计算相邻句子之间的相似度 ----
similarities = []
for i in range(len(sentences) - 1):
sim = cosine_similarity(embeddings[i], embeddings[i+1])
similarities.append(sim)
# 记录下相似度,便于调试
print(f" Sentence {i} -> {i+1}: similarity = {sim:.4f}")
# ---- 4.4 定位分割点 ----
# 分割点放在相似度低于阈值的位置
breakpoints = []
for idx, sim in enumerate(similarities):
if sim < similarity_threshold:
# 分割点位于 idx 和 idx+1 之间
breakpoints.append(idx + 1)
print(f"Detected breakpoints at: {breakpoints}")
# ---- 4.5 按照分割点组合句子生成 chunks ----
chunks = []
start = 0
for bp in breakpoints:
# 如果当前 segment 满足最小句子数要求,则独立成 chunk
if bp - start >= min_chunk_sentences:
chunk_text = " ".join(sentences[start:bp])
chunks.append(chunk_text)
start = bp
# 否则跳过这个分割点,继续向后合并(保证 chunk 不过于零碎)
# (你也可以改为强制分割,取决于业务需求)
# 最后一段剩余句子
if start < len(sentences):
chunk_text = " ".join(sentences[start:])
chunks.append(chunk_text)
return chunks
# ============================================================================
# 5. 演示:用一段多主题的文本测试语义分块
# ============================================================================
if __name__ == "__main__":
# 示例文档包含三个自然段落,话题明显不同
sample_doc = (
"The cat sat on the mat. It was a sunny day. The cat looked very happy. "
"Quantum computing uses qubits instead of classical bits. Qubits can exist in superposition. "
"Entanglement allows qubits to be correlated with each other. "
"The best pasta is made with durum wheat semolina. Fresh pasta requires only eggs and flour. "
"Many Italian grandmothers have their own secret recipe."
)
print("Original document:\n", sample_doc)
print("\n--- Performing Semantic Chunking ---")
result_chunks = semantic_chunk(sample_doc, similarity_threshold=0.75, min_chunk_sentences=2)
print("\n--- Resulting Chunks ---")
for i, chunk in enumerate(result_chunks):
print(f"Chunk {i+1}: {chunk}\n")
Key Takeaways
| 要点 | EN | CN |
|---|---|---|
| 语义分块依据 | Splits are based on semantic similarity, not fixed length. | 切分依据是语义相似度,而非固定长度。 |
| 核心工具 | Embedding model + cosine similarity. | 使用 Embedding 模型 + 余弦相似度。 |
| 分割点判定 | Similarity drops below threshold → new chunk. | 相似度低于阈值 → 分割点。 |
| 最小块约束 | min_chunk_sentences prevents overly small chunks. | 设置最小句子数防止块过小。 |
| 优势 | Keeps complete ideas together, improves retrieval and downstream LLM understanding. | 保持完整语义单元,提高检索和下游 LLM 理解效果。 |
| 生产注意事项 | Use proper sentence tokenizer (nltk/spaCy), handle API rate limits, consider caching embeddings. | 生产中用专业分句工具,注意 API 频率限制,可缓存 embedding。 |

