BM25

What is BM25?

BM25 (Best Matching 25) is a keyword-based ranking algorithm used in information retrieval to score and rank documents based on their relevance to a search query. It’s called “Best Matching 25” because it was the 25th variant in a series of scoring functions proposed by its creators. BM25 is the default ranking algorithm in Elasticsearch and most production search engines

In production RAG systems, BM25 is the standard keyword search algorithm. It has “withstood the test of time for decades since its invention”. Here’s why it matters:

What is it used for?

BM25 is used to find documents that contain the exact words or phrases a user typed into the search box. Its core job is lexical (term-based) retrieval โ€” matching literal text strings between the query and the documents.
BM25 ็”จไบŽๆŸฅๆ‰พๅŒ…ๅซ็”จๆˆท่พ“ๅ…ฅๆœ็ดขๆก†ไธญ็š„็กฎๅˆ‡ๅ•่ฏๆˆ–็Ÿญ่ฏญ็š„ๆ–‡ๆกฃใ€‚ๅฎƒ็š„ๆ ธๅฟƒๅทฅไฝœๆ˜ฏ่ฏๆณ•๏ผˆๅŸบไบŽ่ฏ้กน็š„๏ผ‰ๆฃ€็ดขโ€”โ€”ๅœจๆŸฅ่ฏขๅ’Œๆ–‡ๆกฃไน‹้—ดๅŒน้…ๅญ—้ขไธŠ็š„ๆ–‡ๆœฌๅญ—็ฌฆไธฒใ€‚

ๆƒณ่ฑกไฝ ๆœ‰ไธ€ไธช่ฃ…็€ๆˆๅƒไธŠไธ‡ไปฝๅˆๅŒ็š„ๅทจๅคงๆ–‡ไปถๆŸœใ€‚BM25ๅฐฑๆ˜ฏ้‚ฃไธชๅธฆๆ ‡็ญพ็š„็ดขๅผ•็ณป็ปŸโ€”โ€”ๅฝ“ไฝ ๆœ็ดข”ไธๅฏๆŠ—ๅŠ›ๆกๆฌพ็ฌฌ4.2่Š‚”ๆ—ถ๏ผŒๅฎƒ็›ดๆŽฅ็ฟปๅˆฐๅŒ…ๅซ่ฟ™ไบ›็กฎๅˆ‡ๆ–‡ๅญ—็š„้กต้ขใ€‚ๅฎƒไธ่ฏ•ๅ›พ็†่งฃ”ไธๅฏๆŠ—ๅŠ›”ๆ˜ฏไป€ไนˆๆ„ๆ€๏ผ›ๅฎƒๅชๆ˜ฏๆ‰พๅˆฐ้‚ฃไบ›ๅญ—ๆฏๅ‡บ็Žฐ็š„้กต้ขใ€‚่ฟ™ๅฐฑๆ˜ฏๅฎƒ็š„ๅทฅไฝœ๏ผšๅญ—้ขใ€็ฒพ็กฎใ€ๅฟซ้€Ÿ็š„ๆ–‡ๆœฌๅฎšไฝใ€‚

Specific Use Cases example

ๅœบๆ™ฏ (Scenario)BM25ๅฆ‚ไฝ•ๅ‘ๆŒฅไฝœ็”จ (How BM25 helps)
้”™่ฏฏไปฃ็ ๆŸฅ่ฏข (Error code lookup)User searches “HTTP 503” โ€” BM25 finds the exact doc containing “503”
ไบงๅ“SKU/ๅบๅˆ—ๅท (Product SKU/Serial #)User searches “ABC-123-XYZ” โ€” BM25 precisely matches the alphanumeric string
ไผไธšๅ†…้ƒจๆœฏ่ฏญ (Enterprise jargon)User searches “Databricks Unity Catalog” โ€” BM25 retrieves docs with those specific terms
ๆณ•ๅพ‹/ๅˆ่ง„ๆ–‡ๆกฃ (Legal/Compliance)User searches “GDPR Article 17” โ€” BM25 matches the exact legal reference
ๆ—ฅๅฟ—ๅˆ†ๆž (Log analysis)User searches “TimeoutException at line 42” โ€” BM25 finds the exact log entry

Why use it?

You use BM25 because vector search (dense embeddings) has a fatal weakness: it understands meaning but ignores exact form. Here’s the hard truth:
ไฝ ไฝฟ็”จBM25ๆ˜ฏๅ› ไธบๅ‘้‡ๆฃ€็ดข๏ผˆ็จ ๅฏ†ๅตŒๅ…ฅ๏ผ‰ๆœ‰ไธ€ไธช่‡ดๅ‘ฝๅผฑ็‚น๏ผšๅฎƒ็†่งฃๅซไน‰๏ผŒไฝ†ๅฟฝ็•ฅ็ฒพ็กฎๅฝขๅผใ€‚ไปฅไธ‹ๆ˜ฏ็กฌๆ ธ็œŸ็›ธ๏ผš

Hard Keyword Matching (Serial Numbers, IDs, and Codes)

Vector embeddings compress text into semantic space, which often washes out the exact identity of unique strings like product IDs, error codes (e.g., ERR_404_AUTH), or specific numbers. If a user searches for a specific part number, vector search might return a “similar” product. BM25 treats text as discrete tokens, ensuring exact matches are surfaced instantly.
ๅ‘้‡ๅตŒๅ…ฅๅฐ†ๆ–‡ๆœฌๅŽ‹็ผฉๅˆฐ่ฏญไน‰็ฉบ้—ดไธญ๏ผŒ่ฟ™ๅพ€ๅพ€ไผšๆจก็ณŠๆމๅ”ฏไธ€ๅญ—็ฌฆไธฒ๏ผˆๅฆ‚ไบงๅ“ IDใ€้”™่ฏฏไปฃ็ ๅฆ‚ ERR_404_AUTH ๆˆ–็‰นๅฎšๆ•ฐๅญ—๏ผ‰็š„็ฒพ็กฎ็‰นๅพใ€‚ๅฆ‚ๆžœ็”จๆˆทๆœ็ดข็‰นๅฎš็š„้›ถไปถๅท๏ผŒๅ‘้‡ๆœ็ดขๅฏ่ƒฝไผš่ฟ”ๅ›žไธ€ไธชโ€œ็›ธไผผโ€็š„ไบงๅ“ใ€‚่€Œ BM25 ๅฐ†ๆ–‡ๆœฌ่ง†ไธบ็ฆปๆ•ฃ็š„ Token๏ผŒ่ƒฝ็กฎไฟ็ซ‹ๅณๅŒน้…ๅˆฐ็ฒพ็กฎ็š„็›ฎๆ ‡ใ€‚

The example Scenario

Our database has two documents:

  • Document A: “Troubleshooting guide for pump model P-400. If you encounter error FX-992, it means the pressure valve is jammed. Clear the debris.” (P-400 ๅž‹ๆฐดๆณตๆ•…้šœๆŽ’้™คๆŒ‡ๅ—ใ€‚ๅฆ‚ๆžœ้‡ๅˆฐ FX-992 ้”™่ฏฏ๏ผŒ่ฏดๆ˜ŽๅŽ‹ๅŠ›้˜€ๅกไฝใ€‚่ฏทๆธ…็†็ขŽ็‰‡ใ€‚)
  • Document B: “Troubleshooting guide for pump model P-401. If you encounter error FX-993, it means the pressure valve is jammed. Clear the debris.” (P-401 ๅž‹ๆฐดๆณตๆ•…้šœๆŽ’้™คๆŒ‡ๅ—ใ€‚ๅฆ‚ๆžœ้‡ๅˆฐ FX-993 ้”™่ฏฏ๏ผŒ่ฏดๆ˜ŽๅŽ‹ๅŠ›้˜€ๅกไฝใ€‚่ฏทๆธ…็†็ขŽ็‰‡ใ€‚)

ๅ‘้‡ๆœ็ดขๆ˜ฏๅฆ‚ไฝ•็œ‹ๅพ…่ฟ™ไธช้—ฎ้ข˜็š„๏ผˆไธบไป€ไนˆๅฎƒไผšๅคฑ่ดฅ๏ผ‰

  • EN: Vector search converts the query and documents into lists of numbers (embeddings) based on their meaning.
    • To the vector model, Document A and Document B mean almost the exact same thing: “A troubleshooting guide for a water pump experiencing a jammed pressure valve.”
    • Because the models compress the text, P-400 vs P-401 and FX-992 vs FX-993 look 99% similar in the mathematical “semantic space”.
    • The Result: Vector search might score Document B higher than Document A just by random mathematical variance. If the RAG system feeds Document B to the LLM, the technician gets the wrong repair instructions for a completely different pump model.
  • CN: ๅ‘้‡ๆœ็ดขๆ นๆฎๅซไน‰ๅฐ†ๆŸฅ่ฏข่ฏๅ’Œๆ–‡ๆกฃ่ฝฌๆขไธบไธ€ไธฒๆ•ฐๅญ—๏ผˆๅตŒๅ…ฅๅ‘้‡๏ผ‰ใ€‚
    • ๅฏนไบŽๅ‘้‡ๆจกๅž‹ๆฅ่ฏด๏ผŒๆ–‡ๆกฃ A ๅ’Œๆ–‡ๆกฃ B ็š„ๅซไน‰ๅ‡ ไนŽๅฎŒๅ…จ็›ธๅŒ๏ผšโ€œๅ…ณไบŽๆฐดๆณตๅŽ‹ๅŠ›้˜€ๅกไฝ็š„ๆ•…้šœๆŽ’้™คๆŒ‡ๅ—ใ€‚โ€
    • ๅ› ไธบๆจกๅž‹ๅŽ‹็ผฉไบ†ๆ–‡ๆœฌ๏ผŒP-400 ไธŽ P-401ใ€FX-992 ไธŽ FX-993 ๅœจๆ•ฐๅญฆ็š„โ€œ่ฏญไน‰็ฉบ้—ดโ€ไธญ็œ‹่ตทๆฅๆœ‰ 99% ็š„็›ธไผผๅบฆใ€‚
    • ็ป“ๆžœ๏ผš ๅ‘้‡ๆœ็ดขๅฏ่ƒฝไผšๅ› ไธบ้šๆœบ็š„ๆ•ฐๅญฆๅๅทฎ๏ผŒ็ป™ๆ–‡ๆกฃ B ๆ‰“ๅ‡บๆฏ”ๆ–‡ๆกฃ A ๆ›ด้ซ˜็š„ๅˆ†ๆ•ฐใ€‚ๅฆ‚ๆžœ RAG ็ณป็ปŸๆŠŠๆ–‡ๆกฃ B ๅ–‚็ป™ไบ†ๅคงๆจกๅž‹๏ผˆLLM๏ผ‰๏ผŒๆŠ€ๆœฏไบบๅ‘˜ๅฐฑไผšๅพ—ๅˆฐๅฎŒๅ…จไธๅŒ็š„ๆฐดๆณตๅž‹ๅท็š„้”™่ฏฏ็ปดไฟฎๆŒ‡ไปคใ€‚

Out-of-Vocabulary (OOV) & Domain-Specific Jargon
ๆœช็™ปๅฝ•่ฏ๏ผˆOOV๏ผ‰ไธŽ่กŒไธšไธ“ไธšๆœฏ่ฏญ

Pre-trained embedding models have a fixed vocabulary. When your production data contains highly specialized enterprise jargon, internal acronyms, or brand-new product names, the vector model won’t understand the semantics and will guess poorly. BM25 doesn’t need to “understand” the word; it calculates frequency ($TF$) and rarity ($IDF$), making it incredibly robust for proprietary data.
้ข„่ฎญ็ปƒ็š„ๅตŒๅ…ฅๆจกๅž‹่ฏ่กจๆ˜ฏๅ›บๅฎš็š„ใ€‚ๅฝ“ๆ‚จ็š„็”Ÿไบงๆ•ฐๆฎๅŒ…ๅซ้ซ˜ๅบฆไธ“ไธšๅŒ–็š„ไผไธšๆœฏ่ฏญใ€ๅ†…้ƒจ็ผฉๅ†™ๆˆ–ๅ…จๆ–ฐๅ‘ๅธƒ็š„ไบงๅ“ๅ็งฐๆ—ถ๏ผŒๅ‘้‡ๆจกๅž‹ๆ— ๆณ•็†่งฃๅ…ถ่ฏญไน‰๏ผŒๅช่ƒฝ่ฟ›่กŒ็ณŸ็ณ•็š„็›ฒ็Œœใ€‚BM25 ไธ้œ€่ฆโ€œ็†่งฃโ€่ฟ™ไธช่ฏ๏ผŒๅฎƒ็›ดๆŽฅ่ฎก็ฎ—่ฏ้ข‘๏ผˆ$TF$๏ผ‰ๅ’Œ็จ€็ผบๅบฆ๏ผˆ$IDF$๏ผ‰๏ผŒ่ฟ™ไฝฟๅพ—ๅฎƒๅœจๅค„็†็งๆœ‰ไธ“ๅฑžๆ•ฐๆฎๆ—ถ่กจ็Žฐๅพ—ๅผ‚ๅธธ้ฒๆฃ’ใ€‚

Outperforming on Short Queries

When users type short, concise queries (e.g., “SQL deadlock fix”), vector models sometimes lack enough context to generate a high-quality embedding vector, leading to diluted results. BM25 shines here because it treats those exact terms as heavy anchors, instantly pulling documents containing those exact keywords.
ๅฝ“็”จๆˆท่พ“ๅ…ฅ็ฎ€็Ÿญใ€็ฒพ็‚ผ็š„ๆŸฅ่ฏข๏ผˆไพ‹ๅฆ‚ โ€œSQL ๆญป้”ไฟฎๅคโ€๏ผ‰ๆ—ถ๏ผŒๅ‘้‡ๆจกๅž‹ๆœ‰ๆ—ถไผšๅ› ไธบ็ผบไน่ถณๅคŸ็š„ไธŠไธ‹ๆ–‡่€Œๆ— ๆณ•็”Ÿๆˆ้ซ˜่ดจ้‡็š„ๅตŒๅ…ฅๅ‘้‡๏ผŒๅฏผ่‡ดๆฃ€็ดข็ป“ๆžœ่ขซ็จ€้‡Šใ€‚BM25 ๅœจ่ฟ™็งๅœบๆ™ฏไธ‹ๅคงๆ”พๅผ‚ๅฝฉ๏ผŒๅ› ไธบๅฎƒๅฐ†่ฟ™ไบ›ๅ…ทไฝ“็š„่ฏ่ง†ไธบๆ ธๅฟƒ้”š็‚น๏ผŒ็žฌ้—ดๆ‹‰ๅ–ๅŒ…ๅซ่ฟ™ไบ›็ฒพ็กฎๅ…ณ้”ฎ่ฏ็š„ๆ–‡ๆกฃใ€‚

Cost, Speed, and Scale

Vector databases require specialized, memory-heavy hardware (RAM/GPUs) to perform Approximate Nearest Neighbor (ANN) searches at scale. BM25 runs on highly optimized, inverted indices (via OpenSearch, Elasticsearch, etc.) that are computationally cheap, lightning-fast, and can handle billions of documents with standard CPU architecture.
ๅ‘้‡ๆ•ฐๆฎๅบ“้œ€่ฆไธ“้—จ็š„้ซ˜ๅ†…ๅญ˜็กฌไปถ๏ผˆRAM/GPU๏ผ‰ๆฅๅœจๅคง่ง„ๆจกๆ•ฐๆฎไธ‹ๆ‰ง่กŒ่ฟ‘ไผผๆœ€่ฟ‘้‚ป๏ผˆANN๏ผ‰ๆœ็ดขใ€‚่€Œ BM25 ่ฟ่กŒๅœจ้ซ˜ๅบฆไผ˜ๅŒ–็š„ๅ€’ๆŽ’็ดขๅผ•ไธŠ๏ผˆ้€š่ฟ‡ OpenSearchใ€Elasticsearch ็ญ‰๏ผ‰๏ผŒ่ฎก็ฎ—ๆˆๆœฌๆžไฝŽ๏ผŒ้€Ÿๅบฆๆžๅฟซ๏ผŒไฝฟ็”จๆ ‡ๅ‡†็š„ CPU ๆžถๆž„ๅณๅฏ่ฝปๆพๅค„็†ๆ•ฐ็™พไบฟๆกๆ–‡ๆกฃใ€‚

Decision Matrix example

ๅœบๆ™ฏ (Scenario)ๅช็”จๅ‘้‡ (Vector Only)ๅช็”จBM25 (BM25 Only)ๆททๅˆ (Hybrid)
็”จๆˆทๆœ็ฝ•่ง้”™่ฏฏ็  0xDEADBEEFโŒ ๅคฑ่ดฅ (้›ถๅฌๅ›ž)โœ… ๅฎŒ็พŽโœ… ๅฎŒ็พŽ
็”จๆˆทๆœ้€š็”จๆฆ‚ๅฟต “cloud data integration”โœ… ๅฅฝ (ๅŒไน‰่ฏๆณ›ๅŒ–)โš ๏ธ ไธ€่ˆฌ (ไป…ๅญ—้ข)โœ… ๆœ€ๅฅฝ
ๅปถ่ฟŸๆ•ๆ„Ÿๅž‹/้ซ˜QPSๅœบๆ™ฏโŒ GPUๆ˜‚่ดตโœ… CPUๆžๅฟซโš ๏ธ ๅฏไผ˜ๅŒ–(ไธค้˜ถๆฎต)
ๅˆ่ง„ๅฎก่ฎก้œ€่ฆ่งฃ้‡Šๆฃ€็ดขๅŽŸๅ› โŒ ไธๅฏ่งฃ้‡Šโœ… ๅฎŒๅ…จๅฏ่งฃ้‡Šโœ… ๅฏ่งฃ้‡Š(BM25้ƒจๅˆ†)
ๆ–ฐไบงๅ“ไปฃๅท (Project Athena) ๅˆšๅ‘ๅธƒโŒ ไธ่ฎค่ฏ†โœ… ๅณๆ—ถๆ”ฏๆŒโœ… ๅณๆ—ถๆ”ฏๆŒ

One-Sentence Summary

ou use BM25 not because it’s “better” than vector search, but because it covers the exact-match failure cases that vector search inherently cannot, while simultaneously saving cost, enabling auditability, and serving as the mandatory lexical leg for Hybrid Search (RRF).

What does BM25 include?

Formula Overview: BM25 generates a relevance score for each document-query pair. The total score is the sum of scores for each query term.
BM25ไธบๆฏไธชๆ–‡ๆกฃ-ๆŸฅ่ฏขๅฏน็”Ÿๆˆไธ€ไธช็›ธๅ…ณๆ€งๅˆ†ๆ•ฐใ€‚ๆ€ปๅˆ†ๆ˜ฏๆฏไธชๆŸฅ่ฏข่ฏ้กนๅˆ†ๆ•ฐ็š„ๆ€ปๅ’Œ

Three Core Improvements over TF-IDF:

ๆ”น่ฟ›ENCN
่ฏ้ข‘้ฅฑๅ’Œ (Term Frequency Saturation)Documents get diminishing returns as they include more instances of a keyword. A document with “pizza” 20 times isn’t twice as relevant as one with 10 times.โ€œ่ฟ‡็ŠนไธๅŠ๏ผŒ็‰ฉๆžๅฟ…ๅ๏ผˆๆˆ–่พน้™…ๆ•ˆ็›Š้€’ๅ‡๏ผ‰ใ€‚็ณป็ปŸๅˆคๆ–ญไธ€็ฏ‡ๆ–‡ๆกฃๆ˜ฏๅฆๅ…ณไบŽโ€œๆŠซ่จโ€๏ผŒ็œ‹็š„ๆ˜ฏโ€œๆ˜ฏๅฆๆœ‰โ€๏ผŒ่€Œไธๆ˜ฏโ€œๆœ‰ๅคšๅฐ‘โ€ใ€‚ไธ€ๆ—ฆๆ–‡ๆกฃๆ˜Ž็กฎๆๅˆฐไบ†โ€œๆŠซ่จโ€๏ผŒ็ณป็ปŸๅฐฑๅŸบๆœฌ่ฎคๅฎšๅฎƒ็›ธๅ…ณไบ†ใ€‚ไน‹ๅŽๅ†้‡ๅคๆๅˆฐ๏ผŒๅฏนโ€œ็›ธๅ…ณๆ€งๅˆ†ๆ•ฐโ€็š„ๆๅ‡ๅพฎไนŽๅ…ถๅพฎ๏ผŒ็”š่‡ณๅ› ไธบๅ †็ Œๅ…ณ้”ฎ่ฏ่€Œๅฏผ่‡ดๆ‰ฃๅˆ†ใ€‚
ๆ–‡ๆกฃ้•ฟๅบฆๅฝ’ไธ€ๅŒ– (Document Length Normalization)Penalties for long documents are also diminishing. TF-IDF penalizes long documents too aggressively; BM25 applies diminishing additional penalties.ๆ–‡็ซ ่ถŠ้•ฟ๏ผŒ่ถŠๅฎนๆ˜“่ขซโ€œ่ฏฏไผคโ€๏ผŒไฝ†BM25ๅญฆไผšไบ†โ€œๆ‰‹ไธ‹็•™ๆƒ…โ€๏ผŒ่€Œไธ”่ถŠ้•ฟ่ถŠ็•™ๆƒ…ใ€‚
ไพ‹ๅฆ‚๏ผš่ถ…้•ฟๆ–‡๏ผˆไปŽ1000ๅญ—ๅ˜10000ๅญ—๏ผ‰๏ผšๅœจBM25็œผ้‡Œ๏ผŒๅขžๅŠ ็š„้‚ฃ9000ๅญ—ๅธฆๆฅ็š„โ€œๅ‡ๅˆ†ๆ•ˆๆžœโ€ๆžๅ…ถๅพฎๅผฑใ€‚ๆƒฉ็ฝšๅบ”่ฏฅๆœ‰ไธชโ€œๅฐ้กถๅ€ผโ€ใ€‚ๆ–‡็ซ ๆฏๅ˜้•ฟไธ€ๆˆช๏ผŒๆƒฉ็ฝšๅŠ›ๅบฆไผš่ถŠๆฅ่ถŠๅฐ๏ผˆๅณ้ขๅค–ๆƒฉ็ฝšๆ˜ฏ้€’ๅ‡็š„๏ผ‰ใ€‚
ๅฏ่ฐƒ่ถ…ๅ‚ๆ•ฐ (Tunable Hyperparameters)Two parameters (k1 and b) let you control the degree of term frequency saturation and document length normalization.ไธคไธชๅ‚ๆ•ฐ๏ผˆk1ๅ’Œb๏ผ‰่ฎฉไฝ ๆŽงๅˆถ่ฏ้ข‘้ฅฑๅ’Œๅบฆๅ’Œๆ–‡ๆกฃ้•ฟๅบฆๅฝ’ไธ€ๅŒ–็š„็จ‹ๅบฆใ€‚
k1ๆŽงๅˆถโ€œ่ฏ้ข‘้ฅฑๅ’Œๅบฆโ€๏ผŒๅ…ณ้”ฎ่ฏๅ‡บ็Žฐๅคšๅฐ‘ๆฌกไน‹ๅŽ๏ผŒๅ†ๅŠ ๅˆ†ๅฐฑๆฒกๅ•ฅ็”จไบ†

Parameter k1

k1 ( 1.2-2.0) : Controls term frequency saturation. Higher k1 means term frequency continues to matter more; lower k1 means saturation happens faster.
โ€œๅ…ณ้”ฎ่ฏๅ‡บ็Žฐๅคšๅฐ‘ๆฌกไน‹ๅŽ๏ผŒๅ†ๅŠ ๅˆ†ๅฐฑๆฒกๅ•ฅ็”จไบ†โ€
ๆžไฝŽๅ€ผ๏ผˆๅฆ‚ k1=0๏ผ‰๏ผšๅช่ฆๆ–‡ๆกฃๅ‡บ็Žฐ่ฟ‡โ€œๆŠซ่จโ€๏ผŒๅˆ†ๆ•ฐๅฐฑๅ›บๅฎšไบ†๏ผŒๅŽ้ขๅ‡บ็Žฐ 100 ๆฌกไนŸไธๅŠ ๅˆ†๏ผˆ่ฟ‡ไบŽๆญปๆฟ๏ผ‰ใ€‚
ไธญ้—ดๅ€ผ๏ผˆ้ป˜่ฎค k1=1.2~2.0๏ผ‰๏ผšๅ‡บ็Žฐ็ฌฌ 1 ๆฌกๅŠ  10 ๅˆ†๏ผŒๅ‡บ็Žฐ็ฌฌ 2 ๆฌกๅŠ  5 ๅˆ†๏ผŒๅ‡บ็Žฐ็ฌฌ 3 ๆฌกๅŠ  2 ๅˆ†โ€ฆโ€ฆๅ‡บ็Žฐ็ฌฌ 10 ๆฌกๆ—ถ๏ผŒๅŠ ๅˆ†ๅ‡ ไนŽไธบ 0๏ผˆ่ฟ™ๅฐฑๆ˜ฏไฝ ็ฌฌไธ€่ฝฎ้—ฎ็š„โ€œๆ”ถ็›Š้€’ๅ‡โ€๏ผ‰ใ€‚
ๆž้ซ˜ๅ€ผ๏ผˆๅฆ‚ k1=10๏ผ‰๏ผšๅ‡บ็Žฐ็ฌฌ 1 ๆฌกๅŠ  1 ๅˆ†๏ผŒๅ‡บ็Žฐ็ฌฌ 10 ๆฌกๅŠ  9 ๅˆ†๏ผŒๅ‡ ไนŽๅ‘ˆ็บฟๆ€งๅขž้•ฟ๏ผˆ่ฟ™ๅฐฑ้€€ๅŒ–ๆˆ่€ๆ—ง็š„ TF-IDF๏ผŒๅฎนๆ˜“่ขซๅ…ณ้”ฎ่ฏๅ †็ ŒไฝœๅผŠ๏ผ‰ใ€‚
e.g.
k1 ๅ†ณๅฎšไฝ ็š„โ€œ้ฅญ้‡โ€ใ€‚ๆญฃๅธธไบบ๏ผˆk1=1.2๏ผ‰ๅƒ 3 ๅ—ๆŠซ่จๅฐฑ้ฅฑไบ†๏ผŒๅ†ไธŠ็ฌฌ 10 ๅ—ไนŸๅƒไธไธ‹ไบ†๏ผˆๅˆ†ๆ•ฐๅฐ้กถ๏ผ‰ใ€‚k1 ่ฐƒๅพ—่ถŠไฝŽ๏ผŒไบบ่ถŠๅฎนๆ˜“้ฅฑ๏ผˆ้ฅฑๅ’Œ่ถŠๅฟซ๏ผ‰๏ผ›่ฐƒๅพ—่ถŠ้ซ˜๏ผŒไบบ่ถŠ่ƒฝๅƒ๏ผˆ้ฅฑๅ’Œ่ถŠๆ…ข๏ผ‰ใ€‚

Parameter “b” 

parameter “b” (้€šๅธธ 0.0-1.0) : Controls document length normalization. b=0 means no length normalization; b=1 means full normalization. Default is typically 0.75.
 ็ฎกโ€œๅ› ไธบๆ–‡็ซ ๅคช้•ฟ่€Œๆ‰ฃๅˆ†ๆ—ถ๏ผŒๆ‰ฃๅพ—ๆœ‰ๅคš็‹ โ€๏ผˆๅณไฝ ็ฌฌไบŒ่ฝฎ้—ฎ็š„โ€œ้•ฟๆ–‡ๆกฃๆƒฉ็ฝšโ€๏ผ‰ใ€‚

ๆžไฝŽๅ€ผ๏ผˆb=0๏ผ‰๏ผšๅฎŒๅ…จไธ็œ‹ๆ–‡็ซ ้•ฟๅบฆใ€‚ไธ€็ฏ‡ 10000 ๅญ—็š„็™พ็ง‘ๅ…จไนฆๅ’Œไธ€็ฏ‡ 100 ๅญ—็š„ๅพฎๅš็Ÿญๆ–‡๏ผŒๅ—ๅˆฐ็š„ๅพ…้‡ๅฎŒๅ…จไธ€ๆ ทใ€‚ๅช่ฆ้•ฟๆ–‡้‡Œโ€œๆŠซ่จโ€ๅ‡บ็Žฐๅพ—ๅคš๏ผŒๅฎƒๅฐฑๆฐธ่ฟœๆŽ’็ฌฌไธ€๏ผˆ่ฟ™ไธๅ…ฌๅนณ๏ผ‰ใ€‚
ๆž้ซ˜ๅ€ผ๏ผˆb=1๏ผ‰๏ผšไธฅๆ ผๆŒ‰็…ง้•ฟๅบฆๆฏ”ไพ‹ๆ‰ฃๅˆ†ใ€‚1000 ๅญ—็š„ๆ–‡็ซ ๏ผŒๆƒฉ็ฝšๅฐฑๆ˜ฏ 100 ๅญ—็š„ 10 ๅ€๏ผˆ่ฟ‡ไบŽๆฟ€่ฟ›๏ผŒ่ฟ™ๅฐฑๆ˜ฏ TF-IDF ็š„ๅผŠ็ซฏ๏ผ‰ใ€‚
ไธญ้—ดๅ€ผ๏ผˆ้ป˜่ฎค b=0.75๏ผ‰๏ผšๆŠ˜ไธญๆ–นๆกˆใ€‚้•ฟๆ–‡็ซ ไผš่ขซๆ‰ฃไธ€็‚นๅˆ†๏ผŒไฝ†ๆ‰ฃๅพ—ๅพˆโ€œๆธฉๆŸ”โ€ใ€‚ๅ“ชๆ€•ๆ–‡็ซ ๆœ‰ 10000 ๅญ—๏ผŒๅช่ฆๆ ธๅฟƒๅ‰ๅ‡ ๆฎตๅๅคๆๅˆฐไบ†โ€œๆŠซ่จโ€๏ผŒ็ฎ—ๆณ•ๅฐฑ็Ÿฅ้“ไฝ ็กฎๅฎžๆ˜ฏ่ฎฒๆŠซ่จ็š„๏ผŒไธไผšๅ› ไธบๅŽ้ขๅบŸ่ฏๅคšๅฐฑๆŠŠไฝ ๅฝปๅบ•ๅŸ‹ๆฒกใ€‚
็ฎ€ๅ•็ฒ—ๆšด็š„็†่งฃ๏ผšb ๅ†ณๅฎšโ€œ่ฟ่ดน้™ฉ็š„ๆ‰ฃ่ดนๆ ‡ๅ‡†โ€ใ€‚b=0 ๆ„ๅ‘ณ็€ไนฐ 1 ๆ–คๅ’Œไนฐ 100 ๆ–ค่ฟ่ดนไธ€ๆ ท๏ผˆๅฏน้•ฟๆ–‡ๅคชๅฎฝๅฎน๏ผ‰๏ผ›b=1 ๆ„ๅ‘ณ็€ไนฐ 100 ๆ–ค่ฟ่ดนๆ˜ฏ 1 ๆ–ค็š„ 100 ๅ€๏ผˆๅฏน้•ฟๆ–‡ๅคช่‹›ๅˆป๏ผ‰๏ผ›b=0.75 ๆ„ๅ‘ณ็€ไนฐ 100 ๆ–คๅชๆ”ถ 1.5 ๅ€็š„่ฟ่ดน๏ผŒ่ถŠ้‡ๅŠ ไปท่ถŠๅฐ‘๏ผˆ้€’ๅ‡ๆƒฉ็ฝš๏ผ‰ใ€‚

Code Implementation

Step 0๏ผšEnvironment Setup

pip install rank_bm25
# ๅฆ‚ๆžœไฝ ๅค„็†ไธญๆ–‡๏ผŒๅฎ‰่ฃ…jieba๏ผ›ๅค„็†่‹ฑๆ–‡ๆŽจ่nltk๏ผˆไฝ†ไปฅไธ‹ไปฃ็ ่‡ชๅธฆๆญฃๅˆ™๏ผŒๅฏไธ่ฃ…๏ผ‰
# If handling Chinese, install jieba; for English, nltk is optional (our regex works fine)
pip install jieba numpy

Step1๏ผšDefine Tokenizer โ€” The Most Critical Step

ๅฎšไน‰ๅˆ†่ฏๅ™จ

The tokenizer splits text into “terms”. BM25’s effectiveness depends entirely on this. Never use raw characters. Always: lowercase, strip punctuation, handle mixed languages.
ๅˆ†่ฏๅ™จๅฐ†ๆ–‡ๆœฌๆ‹†ๅˆ†ๆˆ”่ฏ้กน”ใ€‚BM25็š„ๆ•ˆๆžœๅฎŒๅ…จๅ–ๅ†ณไบŽๆญคใ€‚ๆฐธ่ฟœไธ่ฆไฝฟ็”จๅŽŸๅง‹ๅญ—็ฌฆใ€‚ๆ€ปๆ˜ฏ๏ผšๅฐๅ†™ๅŒ–ใ€ๅŽป้™คๆ ‡็‚นใ€ๅค„็†ๆททๅˆ่ฏญ่จ€ใ€‚

# tokenizer_factory.py
# ็›ฎ็š„: ๅˆ›ๅปบ้€‚็”จไบŽBM25็š„ๅฅๅฃฎๅˆ†่ฏๅ™จ
# Purpose: Create a robust tokenizer for BM25

import re
from typing import List

def create_bm25_tokenizer(language: str = "mixed"):
    """
    ๅˆ›ๅปบBM25ไธ“็”จๅˆ†่ฏๅ™จ
    Create a tokenizer specifically for BM25
    
    BM25ๅˆ†่ฏๅ™จ็š„ไธ‰ๅคงๅŽŸๅˆ™ (Three principles for BM25 tokenizer):
    1. ็ปŸไธ€ๅฐๅ†™ (Unified lowercase) โ€” ็กฎไฟ "Timeout" ๅ’Œ "timeout" ่ขซ่ฏ†ๅˆซไธบๅŒไธ€ไธช่ฏ
    2. ๅŽป้™คๆ ‡็‚น (Strip punctuation) โ€” "timeout!" ๅ˜ๆˆ "timeout"
    3. ไฟ็•™ๅญ—ๆฏๆ•ฐๅญ— (Keep alphanumeric) โ€” "ABC-123" ไธญ็š„ "ABC" ๅ’Œ "123" ่ขซไฟ็•™
    """
    
    if language == "zh" or language == "mixed":
        try:
            import jieba
            def tokenizer(text: str) -> List[str]:
                # 1. ็ปŸไธ€ๅฐๅ†™ (Unify case)
                text_lower = text.lower()
                
                # 2. ไฝฟ็”จjiebaๅˆ†่ฏ (Use jieba tokenization)
                # jieba่ƒฝๅคŸๆ™บ่ƒฝๅค„็†ไธญ่‹ฑๆททๅˆๆ–‡ๆœฌ
                # jieba intelligently handles mixed Chinese-English text
                tokens = list(jieba.cut(text_lower))
                
                # 3. ่ฟ‡ๆปคๆމ็บฏ็ฉบ็™ฝๅ’Œๅ•ๅญ—็ฌฆๆ ‡็‚น (Filter out pure whitespace and single-char punctuation)
                # ้‡็‚น: ไฟ็•™ๆœ‰ๆ„ไน‰็š„่ฏ้กน๏ผŒๅŽป้™คๅ™ชๅฃฐ
                # Key point: Keep meaningful tokens, remove noise
                filtered = [t for t in tokens if t.strip() and not re.match(r'^[\W_]+$', t)]
                
                # ๅฆ‚ๆžœๆฒกๆœ‰่ฏ้กน๏ผŒ่ฟ”ๅ›ž็ฉบๅˆ—่กจ๏ผˆๅŽ็ปญไผšๅค„็†๏ผ‰
                return filtered
            return tokenizer
        except ImportError:
            print("โš ๏ธ jiebaๆœชๅฎ‰่ฃ…๏ผŒไฝฟ็”จๅ›ž้€€็š„้€š็”จๅˆ†่ฏๅ™จ")
            # ๅ›ž้€€: ๆญฃๅˆ™ๅŒน้…ๆ‰€ๆœ‰ๅญ—ๆฏๆ•ฐๅญ—ๅบๅˆ— (Fallback: regex match all alphanumeric sequences)
            return _fallback_tokenizer
    else:
        # ็บฏ่‹ฑๆ–‡ๅˆ†่ฏๅ™จ (Pure English tokenizer)
        return _english_tokenizer


def _english_tokenizer(text: str) -> List[str]:
    """่‹ฑๆ–‡ๅˆ†่ฏๅ™จ: ๅฐๅ†™ + ๆญฃๅˆ™ๆๅ–ๅ•่ฏ (Lowercase + regex extract words)"""
    text_lower = text.lower()
    # \b[a-zA-Z0-9_]+\b ๅŒน้…ๆ‰€ๆœ‰ๅ•่ฏ๏ผŒๅŒ…ๆ‹ฌๅธฆไธ‹ๅˆ’็บฟ็š„
    # ้‡็‚น: ่ฟ™ไผšๆŠŠ "error_code" ไฝœไธบไธ€ไธชๆ•ดไฝ“ไฟ็•™๏ผŒ่€Œไธๆ˜ฏๆ‹†ๆˆ "error" ๅ’Œ "code"
    # Key point: This keeps "error_code" as one token, not split into "error" and "code"
    return re.findall(r'\b[a-zA-Z0-9_]+\b', text_lower)


def _fallback_tokenizer(text: str) -> List[str]:
    """ๅ›ž้€€ๅˆ†่ฏๅ™จ: ็บฏๅญ—็ฌฆ็บงๅˆซๆ‹†่งฃ + ่ฟ‡ๆปค (Fallback: character-level split + filter)"""
    text_lower = text.lower()
    # ๅชไฟ็•™ไธญ่‹ฑๆ–‡ๅ’Œๆ•ฐๅญ—ๅญ—็ฌฆ (Keep only Chinese, English letters, and digits)
    chars = [ch for ch in text_lower if ch.isalnum() or ('\u4e00' <= ch <= '\u9fff')]
    # ๆŒ‰็ฉบๆ ผๆˆ–่ฟž็ปญ่‹ฑๆ–‡/ไธญๆ–‡ๅˆ†็ป„๏ผˆ็ฎ€ๅŒ–็‰ˆ๏ผ‰
    # ่ฟ™้‡Œ็ฎ€ๅ•่ฟ”ๅ›žๅญ—็ฌฆๅˆ—่กจ๏ผŒไฝ†็”Ÿไบง็ŽฏๅขƒไธๆŽจ่
    return chars if chars else ["[EMPTY]"]

Step 2๏ผšIndex Building โ€” “Feed” documents to BM25

# build_bm25_index.py
# ็›ฎ็š„: ๅฐ†่ฏญๆ–™ๅบ“ๅˆ†่ฏๅนถๆž„ๅปบBM25็ดขๅผ•
# Purpose: Tokenize corpus and build BM25 index

from rank_bm25 import BM25Okapi
from typing import List

def build_bm25_index(
    corpus: List[str], 
    tokenizer, 
    k1: float = 1.5, 
    b: float = 0.75
) -> tuple[BM25Okapi, List[List[str]]]:
    """
    ๆž„ๅปบBM25็ดขๅผ•
    Build BM25 index
    
    Args:
        corpus: ๅŽŸๅง‹ๆ–‡ๆกฃๅˆ—่กจ (Raw document list)
        tokenizer: ๅˆ†่ฏๅ™จๅ‡ฝๆ•ฐ (Tokenizer function)
        k1: ่ฏ้ข‘้ฅฑๅ’Œๅบฆๅ‚ๆ•ฐ (Term frequency saturation)
        b: ้•ฟๅบฆๅฝ’ไธ€ๅŒ–ๅ‚ๆ•ฐ (Length normalization)
        
    Returns:
        (bm25_object, tokenized_corpus)
    """
    
    print("๐Ÿ”„ ๅผ€ๅง‹ๆž„ๅปบBM25็ดขๅผ• (Building BM25 index)...")
    
    # ============================================================
    # ๆญฅ้ชค2.1: ๅฏน่ฏญๆ–™ๅบ“้€็ฏ‡ๅˆ†่ฏ (Tokenize each document)
    # ============================================================
    tokenized_corpus = []
    empty_doc_count = 0
    
    for idx, doc in enumerate(corpus):
        tokens = tokenizer(doc)
        
        # ๅ…ณ้”ฎๆฃ€ๆŸฅ: ๅฆ‚ๆžœไธ€็ฏ‡ๆ–‡ๆกฃๅˆ†่ฏๅŽๆฒกๆœ‰่ฏ้กน๏ผŒBM25ไผšๆŠฅ้”™ๆˆ–ๅฟฝ็•ฅๅฎƒ
        # Critical check: If a document has no tokens after tokenization, BM25 errors or ignores it
        if not tokens:
            empty_doc_count += 1
            # ๆ’ๅ…ฅไธ€ไธชๅ ไฝ็ฌฆ๏ผŒ็กฎไฟ็ดขๅผ•ๅฏไปฅๆญฃๅธธๅทฅไฝœ
            # Insert a placeholder to keep the index working
            tokens = ["[EMPTY_DOC]"]
        
        tokenized_corpus.append(tokens)
        
        # ่ฐƒ่ฏ•: ๆ‰“ๅฐๅ‰3็ฏ‡ๆ–‡ๆกฃ็š„ๅˆ†่ฏ็ป“ๆžœ (Debug: print first 3)
        if idx < 3:
            print(f"   Doc {idx} tokens: {tokens[:10]}{'...' if len(tokens) > 10 else ''}")
    
    if empty_doc_count > 0:
        print(f"โš ๏ธ  ๆœ‰ {empty_doc_count} ็ฏ‡็ฉบๆ–‡ๆกฃ๏ผŒๅทฒๆ’ๅ…ฅๅ ไฝ็ฌฆ")
    
    # ============================================================
    # ๆญฅ้ชค2.2: ๅˆๅง‹ๅŒ–BM25Okapi (Initialize BM25Okapi)
    # ============================================================
    # ้‡็‚น: BM25Okapi็š„ๆž„้€ ๅ‡ฝๆ•ฐๆŽฅๅ— tokenized corpus, k1, b
    # Key point: BM25Okapi constructor accepts tokenized corpus, k1, b
    bm25 = BM25Okapi(tokenized_corpus, k1=k1, b=b)
    
    print(f"โœ… ็ดขๅผ•ๆž„ๅปบๅฎŒๆˆ (Index built):")
    print(f"   - ๆ–‡ๆกฃๆ•ฐ (Docs): {len(corpus)}")
    print(f"   - ๅนณๅ‡ๆ–‡ๆกฃ้•ฟๅบฆ (AvgDL): {bm25.avgdl:.2f}")
    print(f"   - ๅ‚ๆ•ฐ k1: {k1}, b: {b}")
    
    return bm25, tokenized_corpus

Step3๏ผšExecute Query โ€” Core Retrieval Logic

# query_bm25.py
# ็›ฎ็š„: ๅฏนๆŸฅ่ฏข่ฟ›่กŒๅˆ†่ฏๅนถ่Žทๅ–Top-K็ป“ๆžœ
# Purpose: Tokenize query and retrieve Top-K results

import numpy as np
from typing import List, Tuple

def search_bm25(
    bm25: BM25Okapi,
    tokenizer,
    query: str,
    corpus: List[str],
    top_k: int = 5,
    normalize: bool = True
) -> List[Tuple[str, float]]:
    """
    ๆ‰ง่กŒBM25ๆฃ€็ดข
    Execute BM25 search
    
    ๅฎŒๆ•ดๆต็จ‹:
    1. ๅˆ†่ฏๆŸฅ่ฏข (Tokenize query)
    2. ่ฎก็ฎ—ๆ‰€ๆœ‰ๆ–‡ๆกฃ็š„BM25ๅˆ†ๆ•ฐ (Compute BM25 scores for all docs)
    3. ๆŒ‰ๅˆ†ๆ•ฐ้™ๅบๆŽ’ๅบๅ–top_k (Sort descending and take top_k)
    4. (ๅฏ้€‰) ๅฝ’ไธ€ๅŒ–ๅˆฐ0-1 (Optionally normalize to 0-1)
    """
    
    # ============================================================
    # ๆญฅ้ชค3.1: ๅˆ†่ฏๆŸฅ่ฏข (Tokenize query)
    # ============================================================
    # ้‡็‚น: ๆŸฅ่ฏขๅฟ…้กปไฝฟ็”จไธŽๆ–‡ๆกฃๅฎŒๅ…จ็›ธๅŒ็š„ๅˆ†่ฏๅ™จ
    # Key point: Query MUST use the exact same tokenizer as documents
    query_tokens = tokenizer(query)
    
    # ๅฆ‚ๆžœๆŸฅ่ฏขๅˆ†่ฏๅŽไธบ็ฉบ๏ผŒ่ฟ”ๅ›ž็ฉบ็ป“ๆžœ (If query has no tokens, return empty)
    if not query_tokens:
        print("โš ๏ธ  ๆŸฅ่ฏขๆ— ๆœ‰ๆ•ˆ่ฏ้กน๏ผŒ่ฟ”ๅ›ž็ฉบ็ป“ๆžœ")
        return []
    
    print(f"๐Ÿ” ๆŸฅ่ฏข่ฏ้กน (Query tokens): {query_tokens}")
    
    # ============================================================
    # ๆญฅ้ชค3.2: ่Žทๅ–ๆ‰€ๆœ‰ๆ–‡ๆกฃ็š„ๅŽŸๅง‹BM25ๅˆ†ๆ•ฐ (Get raw BM25 scores)
    # ============================================================
    # get_scores() ่ฟ”ๅ›žไธ€ไธชlist๏ผŒ็ดขๅผ•ๅฏนๅบ”ๆ–‡ๆกฃ้กบๅบ
    # get_scores() returns a list, index aligns with document order
    raw_scores = bm25.get_scores(query_tokens)
    
    # ============================================================
    # ๆญฅ้ชค3.3: ๆŽ’ๅบๅนถๆๅ–Top-K (Sort and extract Top-K)
    # ============================================================
    # ๅฐ†ๅˆ†ๆ•ฐๅ’Œ็ดขๅผ•้…ๅฏน (Pair scores with indices)
    scored_docs = [(idx, score) for idx, score in enumerate(raw_scores)]
    
    # ๆŒ‰ๅˆ†ๆ•ฐ้™ๅบๆŽ’ๅบ (Sort by score descending)
    # ้‡็‚น: ไฝฟ็”จkey=lambda x: x[1] ๆŒ‰ๅˆ†ๆ•ฐๆŽ’ๅบ
    # Key point: Use key=lambda x: x[1] to sort by score
    sorted_docs = sorted(scored_docs, key=lambda x: x[1], reverse=True)
    
    # ่ฟ‡ๆปคๆމๅˆ†ๆ•ฐไธบ0็š„็ป“ๆžœ๏ผˆๅฏ้€‰๏ผ‰ (Filter out zero-score results โ€” optional)
    # ๆ„ไน‰: ๅˆ†ๆ•ฐไธบ0ๆ„ๅ‘ณ็€ๆฒกๆœ‰ไปปไฝ•ๆŸฅ่ฏข่ฏ้กนๅ‡บ็Žฐๅœจ่ฏฅๆ–‡ๆกฃไธญ๏ผŒๅฎŒๅ…จๆ— ๅ…ณ
    # Meaning: Score 0 means no query term appears in the document; completely irrelevant
    relevant_docs = [(idx, score) for idx, score in sorted_docs if score > 0]
    
    # ๅ–ๅ‰top_kไธช (Take top_k)
    top_k_docs = relevant_docs[:top_k]
    
    # ============================================================
    # ๆญฅ้ชค3.4: ๅฝ’ไธ€ๅŒ–ๅˆ†ๆ•ฐ (Normalize scores โ€” optional)
    # ============================================================
    # ็›ฎ็š„: ๅฐ†ๅˆ†ๆ•ฐๆ˜ ๅฐ„ๅˆฐ0-1ไน‹้—ด๏ผŒไพฟไบŽไบบ็ฑป้˜…่ฏปๅ’ŒๅŽ็ปญ่žๅˆ(RRFไธ้œ€่ฆ)
    # Purpose: Map scores to 0-1 for readability (RRF doesn't need this)
    if normalize and top_k_docs:
        max_score = top_k_docs[0][1]  # ๆœ€้ซ˜ๅˆ† (Highest score)
        if max_score > 0:
            normalized_results = [
                (corpus[idx], score / max_score) 
                for idx, score in top_k_docs
            ]
            return normalized_results
    
    # ่ฟ”ๅ›žๅŽŸๅง‹ๅˆ†ๆ•ฐ (Return raw scores)
    return [(corpus[idx], score) for idx, score in top_k_docs]

End-to-End Complete Example โ€” Put It All Together

# bm25_complete_pipeline.py
# ็›ฎ็š„: ๅฎŒๆ•ด็š„BM25ๆฃ€็ดขๆตๆฐด็บฟ๏ผˆไปŽๅŽŸๅง‹ๆ–‡ๆœฌๅˆฐๆฃ€็ดข็ป“ๆžœ๏ผ‰
# Purpose: Complete BM25 retrieval pipeline (from raw text to retrieval results)

from rank_bm25 import BM25Okapi
import numpy as np
import re
from typing import List, Tuple

# ============================================================
# 1. ๅฎšไน‰ๅˆ†่ฏๅ™จ (Define Tokenizer)
# ============================================================
def simple_mixed_tokenizer(text: str) -> List[str]:
    """
    ๆททๅˆ่ฏญ่จ€ๅˆ†่ฏๅ™จ๏ผˆๆ— ้œ€้ขๅค–ๅบ“๏ผ‰
    Mixed language tokenizer (no extra libraries required)
    
    ๅค„็†ๆ–นๆณ•:
    1. ็ปŸไธ€ๅฐๅ†™ (Lowercase)
    2. ็”จๆญฃๅˆ™ๆๅ–ๆ‰€ๆœ‰ๅญ—ๆฏๆ•ฐๅญ—ๅบๅˆ— (Extract all alphanumeric sequences)
    3. ๅŒๆ—ถไฟ็•™ไธญๆ–‡ๅญ—็ฌฆ (Keep Chinese characters too)
    """
    text_lower = text.lower()
    # ๅŒน้…: ่‹ฑๆ–‡ๅญ—ๆฏ+ๆ•ฐๅญ—+ไธ‹ๅˆ’็บฟ๏ผŒไปฅๅŠไธญๆ–‡ๅญ—็ฌฆ
    # Match: English letters + digits + underscores, and Chinese characters
    # ้‡็‚น: ่ฟ™่ƒฝๅค„็† "Azureๆ•ฐๆฎๅทฅๅŽ‚" ่ฟ™ๆ ท็š„ๆททๅˆๆ–‡ๆœฌ
    # Key point: This handles mixed text like "Azureๆ•ฐๆฎๅทฅๅŽ‚"
    pattern = r'[a-zA-Z0-9_]+|[\u4e00-\u9fa5]+'
    tokens = re.findall(pattern, text_lower)
    
    # ่ฟ‡ๆปคๆމ็บฏๆ•ฐๅญ—๏ผˆๅฏ้€‰๏ผŒ่ง†ๆƒ…ๅ†ต่€Œๅฎš๏ผ‰
    # ๅฆ‚ๆžœไฝ ็š„ๅœบๆ™ฏ้œ€่ฆๅŒน้…ๆ•ฐๅญ—๏ผˆๅฆ‚้”™่ฏฏ็ 500๏ผ‰๏ผŒไธ่ฆ่ฟ‡ๆปค๏ผ
    # ่ฟ™้‡Œไฟ็•™ๆ‰€ๆœ‰token
    # Filter out pure numbers (optional). If you need error codes like 500, DON'T filter!
    # We'll keep all tokens here.
    
    return tokens if tokens else ["[EMPTY]"]


# ============================================================
# 2. ๅ‡†ๅค‡ๆ•ฐๆฎ (Prepare Data)
# ============================================================
corpus = [
    "Azure Data Factory is a cloud-based data integration service",
    "Databricks is a unified analytics platform for data engineering and ML",
    "BM25 is a probabilistic ranking algorithm used in information retrieval",
    "Azure AI Foundry provides tools for building enterprise AI applications",
    "When Azure Function App times out, error code 500 is returned",
    "Databricks Photon engine accelerates query performance on large datasets",
    "The data pipeline uses Event Hubs to ingest streaming data",
    "Error 500: Internal Server Error - check the application logs",
]

# ============================================================
# 3. ๆž„ๅปบ็ดขๅผ• (Build Index)
# ============================================================
print("="*70)
print("ๆญฅ้ชค 1: ๅˆ†่ฏๅนถๆž„ๅปบ็ดขๅผ• (Tokenize & Build Index)")
print("="*70)

tokenized_corpus = []
for doc in corpus:
    tokens = simple_mixed_tokenizer(doc)
    tokenized_corpus.append(tokens)

# ๆ‰“ๅฐๅˆ†่ฏ็ป“ๆžœ้ข„่งˆ (Print tokenization preview)
for i, tokens in enumerate(tokenized_corpus[:3]):
    print(f"Doc {i}: {tokens}")

# ๅˆๅง‹ๅŒ–BM25 (Initialize BM25)
# ไฝฟ็”จ้ป˜่ฎคๅ‚ๆ•ฐ k1=1.5, b=0.75
bm25 = BM25Okapi(tokenized_corpus, k1=1.5, b=0.75)

print(f"\nโœ… ็ดขๅผ•ๆž„ๅปบๅฎŒๆˆ (Index built)")
print(f"   ๆ–‡ๆกฃๆ•ฐ (Docs): {len(corpus)}")
print(f"   ๅนณๅ‡ๆ–‡ๆกฃ้•ฟๅบฆ (AvgDL): {bm25.avgdl:.2f}")

# ============================================================
# 4. ๆ‰ง่กŒๆŸฅ่ฏข (Execute Query)
# ============================================================
print("\n" + "="*70)
print("ๆญฅ้ชค 2: ๆ‰ง่กŒๆŸฅ่ฏข (Execute Query)")
print("="*70)

# ๆŸฅ่ฏข1: ็ฒพ็กฎ้”™่ฏฏ็  (Exact error code)
query1 = "error code 500"
query1_tokens = simple_mixed_tokenizer(query1)
print(f"ๆŸฅ่ฏข1่ฏ้กน: {query1_tokens}")

scores1 = bm25.get_scores(query1_tokens)
# ่Žทๅ–ๅ‰3ไธช็ป“ๆžœ (Get top 3)
top_indices1 = np.argsort(scores1)[::-1][:3]

print(f"\n๐Ÿ” ๆŸฅ่ฏข (Query): '{query1}'")
print("็ป“ๆžœ (Results):")
for rank, idx in enumerate(top_indices1):
    if scores1[idx] > 0:
        print(f"  {rank+1}. ๅˆ†ๆ•ฐ {scores1[idx]:.4f} -> {corpus[idx]}")

# ๆŸฅ่ฏข2: ไบงๅ“ๅ็งฐ + ๅŠŸ่ƒฝ (Product name + feature)
query2 = "Databricks Photon acceleration"
query2_tokens = simple_mixed_tokenizer(query2)
print(f"\nๆŸฅ่ฏข2่ฏ้กน: {query2_tokens}")

scores2 = bm25.get_scores(query2_tokens)
top_indices2 = np.argsort(scores2)[::-1][:3]

print(f"\n๐Ÿ” ๆŸฅ่ฏข (Query): '{query2}'")
print("็ป“ๆžœ (Results):")
for rank, idx in enumerate(top_indices2):
    if scores2[idx] > 0:
        print(f"  {rank+1}. ๅˆ†ๆ•ฐ {scores2[idx]:.4f} -> {corpus[idx]}")

# ๆŸฅ่ฏข3: ๆททๅˆไธญ่‹ฑ (Mixed CN-EN)
query3 = "ๆ•ฐๆฎๅทฅ็จ‹ pipeline ่ถ…ๆ—ถ"
query3_tokens = simple_mixed_tokenizer(query3)
print(f"\nๆŸฅ่ฏข3่ฏ้กน: {query3_tokens}")

scores3 = bm25.get_scores(query3_tokens)
top_indices3 = np.argsort(scores3)[::-1][:3]

print(f"\n๐Ÿ” ๆŸฅ่ฏข (Query): '{query3}'")
print("็ป“ๆžœ (Results):")
for rank, idx in enumerate(top_indices3):
    if scores3[idx] > 0:
        print(f"  {rank+1}. ๅˆ†ๆ•ฐ {scores3[idx]:.4f} -> {corpus[idx]}")

# ============================================================
# 5. ๅ‚ๆ•ฐ่ฐƒไผ˜ๆผ”็คบ (Parameter Tuning Demo)
# ============================================================
print("\n" + "="*70)
print("ๆญฅ้ชค 3: ๅ‚ๆ•ฐ่ฐƒไผ˜ (Parameter Tuning)")
print("="*70)

# ๅฐ่ฏ•ไธๅŒ็š„k1ๅ€ผ (Try different k1 values)
test_k1s = [1.0, 1.5, 2.0]
test_query = "Azure Function App timeout"

print(f"ๆต‹่ฏ•ๆŸฅ่ฏข: '{test_query}'")
print("ไธๅŒk1ๅ€ผๅฏน็ป“ๆžœ็š„ๅฝฑๅ“ (Impact of different k1 values):")

for k1 in test_k1s:
    # ้‡ๆ–ฐๅˆ›ๅปบBM25ๅฏน่ฑก (Re-create BM25 object)
    temp_bm25 = BM25Okapi(tokenized_corpus, k1=k1, b=0.75)
    tokens = simple_mixed_tokenizer(test_query)
    scores = temp_bm25.get_scores(tokens)
    top_idx = np.argmax(scores)  # ๅ–ๆœ€้ซ˜ๅˆ†ๆ–‡ๆกฃ (Take highest scoring doc)
    print(f"  k1={k1:.1f}: ๆœ€ไฝณๆ–‡ๆกฃ็ดขๅผ• {top_idx} -> '{corpus[top_idx][:50]}...' (score: {scores[top_idx]:.4f})")

Key Takeaways

่ฆ็‚นENCN
BM25่งฃๅ†ณๅ‘้‡ๆฃ€็ดข็š„”็ฒพ็กฎๅŒน้…็›ฒๅŒบ”้—ฎ้ข˜BM25 solves the “exact match blind spot” of vector retrievalBM25่งฃๅ†ณๅ‘้‡ๆฃ€็ดข็š„”็ฒพ็กฎๅŒน้…็›ฒๅŒบ”้—ฎ้ข˜
BM25ๆ˜ฏๆททๅˆๆฃ€็ดข๏ผˆHybrid Search๏ผ‰็š„ๅ…ณ้”ฎ”ๅ…ณ้”ฎ่ฏๅˆ†ๆ”ฏ”BM25 is the essential “keyword leg” of Hybrid SearchBM25ๆ˜ฏๆททๅˆๆฃ€็ดข๏ผˆHybrid Search๏ผ‰็š„ๅ…ณ้”ฎ”ๅ…ณ้”ฎ่ฏๅˆ†ๆ”ฏ”
k1ๆŽงๅˆถ่ฏ้ข‘้ฅฑๅ’Œๅบฆ๏ผšๅ€ผ่ถŠ้ซ˜๏ผŒ้ซ˜้ข‘่ฏ่ดก็Œฎ่ถŠๅคงk1 controls term frequency saturation: higher value = more contribution from high-frequency termsk1ๆŽงๅˆถ่ฏ้ข‘้ฅฑๅ’Œๅบฆ๏ผšๅ€ผ่ถŠ้ซ˜๏ผŒ้ซ˜้ข‘่ฏ่ดก็Œฎ่ถŠๅคง
bๆŽงๅˆถๆ–‡ๆกฃ้•ฟๅบฆๅฝ’ไธ€ๅŒ–๏ผšb=1ๅฎŒๅ…จๆƒฉ็ฝš้•ฟๆ–‡ๆกฃ๏ผŒb=0ไธๆƒฉ็ฝšb controls document length normalization: b=1 fully penalizes long docs, b=0 doesn’tbๆŽงๅˆถๆ–‡ๆกฃ้•ฟๅบฆๅฝ’ไธ€ๅŒ–๏ผšb=1ๅฎŒๅ…จๆƒฉ็ฝš้•ฟๆ–‡ๆกฃ๏ผŒb=0ไธๆƒฉ็ฝš
ๆŸฅ่ฏขๅ’Œๆ–‡ๆกฃๅฟ…้กปไฝฟ็”จๅฎŒๅ…จ็›ธๅŒ็š„ๅˆ†่ฏๅ’Œ้ข„ๅค„็†ๆต็จ‹Query and documents must use exactly the same tokenization and preprocessing pipelineๆŸฅ่ฏขๅ’Œๆ–‡ๆกฃๅฟ…้กปไฝฟ็”จๅฎŒๅ…จ็›ธๅŒ็š„ๅˆ†่ฏๅ’Œ้ข„ๅค„็†ๆต็จ‹
BM25่ฟ่กŒๅœจCPUไธŠ๏ผŒๅปถ่ฟŸไฝŽ๏ผˆ~50ms๏ผ‰๏ผŒๆˆๆœฌๆžไฝŽBM25 runs on CPU with low latency (~50ms) and extremely low costBM25่ฟ่กŒๅœจCPUไธŠ๏ผŒๅปถ่ฟŸไฝŽ๏ผˆ~50ms๏ผ‰๏ผŒๆˆๆœฌๆžไฝŽ
็”Ÿไบง้ƒจ็ฝฒๆ—ถ๏ผŒ็”จ้ชŒ่ฏ้›†ๅฏนk1ๅ’Œbๅš็ฝ‘ๆ ผๆœ็ดข่ฐƒไผ˜In production, tune k1 and b via grid search on a validation set็”Ÿไบง้ƒจ็ฝฒๆ—ถ๏ผŒ็”จ้ชŒ่ฏ้›†ๅฏนk1ๅ’Œbๅš็ฝ‘ๆ ผๆœ็ดข่ฐƒไผ˜
ไธ‹ไธ€ไธชๅ•ๅ…ƒ๏ผˆB07๏ผ‰ๅฐ†็”จRRF่žๅˆBM25ๅ’Œๅ‘้‡ๆฃ€็ดขThe next unit (B07) will fuse BM25 with vector retrieval using RRFไธ‹ไธ€ไธชๅ•ๅ…ƒ๏ผˆB07๏ผ‰ๅฐ†็”จRRF่žๅˆBM25ๅ’Œๅ‘้‡ๆฃ€็ดข

Azure AI Foundry service

Step 1: Setup Azure AI Foundry

Assuring you have known how to use Azure Portal and add azure service. I will skip the adding Azure OpenAI service.

Once you add Azure OpenAI service, open “Explore Foundry port” to open Foundry dashboard. the alternative uses https://ai.azure.com/

Recommend you switch to new Foundry. You will be asked either select a existed project or create a new project.

Creating new project is easy, simply follow the screen steps. you cannot miss it.

you can see that you are able to create agents, Explore playgrounds and Find modules and you recent done works.

1) Create Agents

Create Agents = Build your own AI assistant
ๅˆ›ๅปบไธ€ไธชโ€œAIๅŠฉๆ‰‹/ๆ™บ่ƒฝไฝ“โ€

You use this when you want to:

  • define a role (e.g. โ€œData Analyst Agentโ€)
  • add instructions (system prompt) / ๅ†™่ง„ๅˆ™
  • connect tools (SQL, API, files) / ๅŠ ๅทฅๅ…ท๏ผˆSQL / API / ๆ–‡ไปถ๏ผ‰
  • add knowledge (RAG) / ๅŠ ็Ÿฅ่ฏ†ๅบ“๏ผˆRAG
  • make it do tasks automatically

๐Ÿ‘‰ Result: a custom AI agent / ้€ ไธ€ไธชAIๅ‘˜ๅทฅ, ไธ€ไธชโ€œ่ƒฝๅนฒๆดป็š„AIโ€

2) Explore Playgrounds

Playgrounds = Testing area for models
็ ‚็ฎฑ็ณป็ปŸ๏ผŒ ๆจกๅž‹่ฏ•้ชŒๅฎค

You use it to:

  • chat with models (GPT, DeepSeek, etc.) / ๆต‹่ฏ•ไธๅŒๆจกๅž‹๏ผŒ GPT๏ผŒ DeepSeek .
  • test prompts / ๅ†™ prompt ็œ‹ๆ•ˆๆžœ
  • try settings (temperature, tokens) / ่ฐƒๅ‚ๆ•ฐ๏ผˆtemperature ็ญ‰๏ผ‰
  • compare responses / ๅšๅฎž้ชŒ

๐Ÿ‘‰ It is NOT production
๐Ÿ‘‰ It is for experimenting
๐Ÿ‘‰ It is a โ€œSandbox / practice roomโ€

3) Find Models

Find Models = Choose AI model

You use it to:

  • browse available models (GPT-4.x, GPT-5.x, DeepSeek, etc.) / ็œ‹ๆœ‰ๅ“ชไบ›ๅฏไปฅ็”จ็š„ๆจกๅž‹๏ผˆGPTใ€DeepSeek็ญ‰๏ผ‰
  • check capabilities / ๅฏนๆฏ”่ƒฝๅŠ›
  • compare cost/performance / ็œ‹ไปทๆ ผ๏ผŒๆฏ”ๆ€ง่ƒฝ
  • decide which model to deploy

๐Ÿ‘‰ It is the โ€œmodel catalogโ€ / ๅฐฑๆ˜ฏโ€œๆŒ‘AIๅคง่„‘โ€

simply think as:

  • Find Models โ†’ choose the brain / ้€‰ๅ‘˜ๅทฅๅ€™้€‰ไบบ
  • Playgrounds โ†’ test the brain / ้ข่ฏ•ๆต‹่ฏ•
  • Create Agents โ†’ build a worker using the brain / ๆญฃๅผ้›‡ไฝฃ + ๅˆ†้…ๅทฅไฝœ

Step 2: Deploy model

From Project dashboard, click “Find models”, you will find many models over there to be selected. e.g. gpt-chat5.4, DeepSeek-V4-Flash, etc.

็›ด็™ฝๅœฐ่ฏดไบบ่ฏ๏ผšๅฎ‰่ฃ…ไธ€ไธชmodelใ€‚

Choose a one you like, then click “Deploy”, “deploy” is done.



Step 3: Create Agent

From Project dashboard, click Create agents

follow steps to create a agent. It is straight forward. no any confusing. Fill in agent Name.

you create the agent. looks this:


Tool = giving the AI external capabilities.
็ป™ AI ๅŠ โ€œๅค–้ƒจ่ƒฝๅŠ›โ€. ๆฒกๆœ‰tools๏ผŒ AIๅช่ƒฝ่Šๅคฉ-Chat๏ผŒ ๆœ‰tools๏ผŒ AIๆ‰่ƒฝๅšไบ‹ใ€‚

From Agent UI, you can see:

What is “Create toolbox”

Create toolbox = create a container/group for tools

Inside toolbox you can later add:

  • APIs
  • Functions
  • Search
  • Database tools
  • Custom tools
What is “Connect a tool”

Connect a tool = connect an actual usable tool/service/API

Examples:

  • Bing Search
  • Azure AI Search
  • Function API
  • REST API
  • SQL
  • OpenAPI service

First – Create toolbox

cleck “Create toolbox”

Second add tools to toolsBox

Then click Add to add tools into the toolBox

Let’s add a “Bing Search” as example.
cleck “Web search” –> Add tools

Add another Tool – Function / REST API, let Agent call external servicers.

{
  "openapi": "3.0.0",
  "info": {
    "title": "Users API",
    "version": "1.0.0"
  },
  "servers": [
    {
      "url": "https://jsonplaceholder.typicode.com"
    }
  ],
  "paths": {
    "/users": {
      "get": {
        "operationId": "getUsers",
        "summary": "Get list of users",
        "responses": {
          "200": {
            "description": "Successful response"
          }
        }
      }
    }
  }
}

now we have added 2 tools

Test the tools

Go to Agent playground: Agent โ†’ Chat / Playground / Test panel

Test 1: “REST API”

Typing “Use the REST API tool to get all users and show them.”

Test 2: Bing Search

Typing “Use Bing Search tool to find latest information about Azure AI Foundry.”

FORCE the Agent to call the tool

Assuming we have 100 REST API endpoints, each one will return different data, such as the userโ€™s name or the companyโ€™s name, sale’s amount ……
When we add each API Endpoint to ToolBox, we have to give clearly, specifically descriptions. Agent will scan description, it will choose the most specific one to call,

In actual AI project, most case is using “Tag”,

e.g.
Tool Registry:
– name
– description
– schema
– tags

Tool: get_company_financials
Tags: finance, company, revenue, kpi

Tool: get_user_profile
Tags: user, identity, profile


Add below “instructure” On the “Agent UI” –> “Instructions”

“You have access to multiple tools.
Each tool has a description that defines its purpose.
Always:
– Read tool descriptions carefully
– Select the most relevant tool based on semantic meaning of the user request
– Do NOT rely on hardcoded routing rules
– If multiple tools are relevant, choose the most specific one”



RAG = AI answers using retrieved documents instead of memory.

AI retrieves real documents first, then generates answer


Documents (PDF / Word / Wiki)
        โ†“
   Chunking (ๅˆ‡ๅ—)
        โ†“
   Embeddings (ๅ‘้‡ๅŒ–)
        โ†“
   Vector Search (็›ธไผผๅบฆๆฃ€็ดข)
        โ†“
   Retrieved Context
        โ†“
   LLM Answer

1: From Agent UI

FoundryIQ is Microsoftโ€™s managed knowledge system for RAG (Retrieval-Augmented Generation) inside Azure AI Foundry.

click “Connect to Foundry IQ”

1. Create a AI Search:

If you have not Create Azure AI Search Resource, or says Create an Azure AI Search service from Azure Portal, this is the 1st step.

โ€œAI Search Resourceโ€ = search engine server
Azure AI Search Index” = searchable dataset inside it

AI Search itself is acting as the Vector Database.

from azure portal –> AI search

after successfully created AI search resource, will see

We can see 3 parts from AI Search dasjboard:

  • Build your knowledge base
  • Connect your data
  • Monitor and scale

Build your knowledge base: Build a RAG-ready knowledge system
Including:

  • document ingestion
  • indexing
  • embeddings
  • retrieval
  • grounded chat playground

Connect your data: This is where you IMPORT your enterprise data.
e.g.

  • Cosmos DB
  • PDF
  • Blob storage
  • SharePoint
  • SQL

This step creates search indexes.

Monitor and scale: Infrastructure management: scaling, replicas, partitions, performance

2. Build your knowledge base

This step we will create knowledge source. Turn your data into an agentic knowledge base.

To Build your knowledge base, from AI Search Service dashboard, click “Build”

click “Create new” to create knowledge source.

Indexed = Azure stores/searches your processed data locally
Remote = Azure queries external systems live at runtime

Let’s use Azure blob (indexed) as example.

3. Enable text vectorization

This creates:

  • embeddings
  • vector fields
  • semantic retrieval capability

save it, then we see this:

Now, we have successfully built:

  • Blob Storage ingestion
  • Azure AI Search indexing
  • Knowledge Base connection
  • Vectorization enabled (semantic search ready)

๐Ÿ‘‰ In short: your RAG data layer is READY.

Attach Knowledge Source to Agent

From Project UI

create a new base in mainri-ai-search

return Agent UI –> click Add (Knowledge) –> connect to Foundry IQ
now, click “Create a new base in Mainri-ai-search”

Knowledge Base (Index creation wizard)

Test the RAG

Since we have upload company’s “return” and “policy” to blob, let’s test. it works. Agent read the company’s policy doc, and used it to answer my question
“What WFH – please answering in both EN and CN”


LLM Fundamentals

Azure AI Foundry is a Microsoft’s unified Azure platform-as-a-service offering for enterprise AI operations, model builders, and application development. 
ๅพฎ่ฝฏๆ–ฐ็š„ไผไธš็บง AI ๅนณๅฐ๏ผŒไธป่ฆ็”จไบŽๅผ€ๅ‘ใ€‚

  • AI apps / AI ๅบ”็”จ
  • Copilots / Copilot
  • AI agents / AI Agent (ๆ™บ่ƒฝ็ณป็ปŸ)
  • RAG systems / RAG ็ณป็ปŸ
  • enterprise AI workflows / ไผไธšๆ™บ่ƒฝๅทฅไฝœๆต

It is becoming Microsoft’s main AI engineering platform. Think of it as
ๅฎƒๆญฃๅœจๅ˜ๆˆๅพฎ่ฝฏไธป่ฆ็š„AIๅทฅ็จ‹ๅนณๅฐ๏ผŒๆœฌ่ดจไธŠๅฏไปฅ็†่งฃๆˆ

Azure AI Foundry = 
  Azure OpenAI
    + Prompt ็ฎก็†
    + AI Orchestration
    + Agent Framework
    + RAG
    + Evaluation
    + Deployment
    + Monitoring

What does it do? It helps companies

  • build GenAI apps / ๆž„ๅปบ AI ็ณป็ปŸ
  • connect enterprise data / ่ฟžๆŽฅไผไธšๆ•ฐๆฎ
  • orchestrate AI workflows
  • RAG / ๅš RAG
  • manage prompts / ็ฎก็† Prompt
  • Mange Agent / ็ฎก็†ๆ™บ่ƒฝ็ณป็ปŸ
  • evaluate AI quality / ็›‘ๆŽง AI ่ดจ้‡
  • deploy AI safely / ้ƒจ็ฝฒ AI

Key Components / ๆ ธๅฟƒ็ป„ๆˆ

A. Model Access / ๆจกๅž‹็ฎก็†

via / ้€š่ฟ‡:

  • Azure OpenAI
  • model catalog

Use models like / ่ฐƒ็”จๆจกๅž‹:

  • GPT-4
  • GPT-4o
  • open-source models
B. Prompt Flow

Visual orchestration for:

  • prompts / ้“พๆŽฅPrompt
  • workflows / ็ป„็ป‡ๅทฅไฝœๆต
  • chaining / ่ฐƒ่ฏ•
  • testing / ๆต‹่ฏ•
C. RAG

Connect AI to:

  • SharePoint
  • PDFs / ๆ–‡ๆกฃ
  • databases / ไผไธšๆ•ฐๆฎๅบ“
  • enterprise documents / ไผไธšๆ–‡ๆกฃ
D. AI Agents

Build agents that can /ๆž„ๅปบๅฏ่‡ชๅŠจๆ‰ง่กŒไปปๅŠก็š„ๆ™บ่ƒฝ็ณป็ปŸ๏ผˆAgent๏ผ‰:

  • use tools / Tool calling
  • call APIs / ่ฐƒ็”จAPI
  • automate workflows / ่‡ชๅŠจๅทฅไฝœๆต
  • reason across tasks / ๆŽจ็†๏ผŒ่‡ชๅŠจๅˆ†ๆž
E. Evaluation & Monitoring
็›‘ๆŽง

Measure:

  • hallucination
  • safety
  • quality
  • groundedness

Enterprise companies care about this heavily / ไผไธšๆžๅ…ถ้‡่ง†่ฟ™ไธช.


An Agent = LLM + Tools + Memory + Planning

Agent can:

  • decide steps / ่‡ชๅŠจๆ‹†่งฃไปปๅŠก
  • call tools (search, DB, API, code)
  • store memory
  • execute workflows

๐Ÿ‘‰ Think:

โ€œYou give goal โ†’ agent figures out how to achieve itโ€


What is BPE?

BPE (Byte-Pair Encoding) is an algorithm that splits text into tokens by repeatedly merging the most frequent adjacent pairs of characters.
ๆ˜ฏไธ€็งๅฐ†ๆ–‡ๆœฌๆ‹†ๅˆ†ๆˆ token ็š„็ฎ—ๆณ•๏ผŒๅฎƒ้€š่ฟ‡ๅๅคๅˆๅนถๆœ€ๅธธๅ‡บ็Žฐ็š„็›ธ้‚ปๅญ—็ฌฆๅฏนๆฅๆž„ๅปบ่ฏๆฑ‡่กจใ€‚

BPE starts with a base vocabulary of bytes/characters and iteratively merges the most frequent adjacent pairs across a large text corpus.

Three key points

CNEN
1. ่พ“ๅ…ฅๆ˜ฏๆ–‡ๆœฌ1. Input is text
2. ่พ“ๅ‡บๆ˜ฏไธ€ๅฅ—ๅˆๅนถ่ง„ๅˆ™ + token ๅบๅˆ—2. Output is a set of merge rules + a token sequence
3. ๆ ธๅฟƒๆ“ไฝœ๏ผšๆ‰พๆœ€้ข‘็น็š„็›ธ้‚ปๅฏน๏ผŒๅˆๅนถ๏ผŒ้‡ๅค3. Core operation: find the most frequent adjacent pair, merge, repeat

BPE ๆฏไธ€ๆญฅๅช็œ‹็›ธ้‚ป็š„ไธคไธชๅญ—็ฌฆ๏ผˆๆˆ–ไธคไธช token๏ผ‰ใ€‚่ฟ™ๅฐฑๆ˜ฏไธบไป€ไนˆๅซ Byte-Pair๏ผˆๅญ—่Š‚ๅฏน๏ผ‰โ€”โ€” ๆฏๆฌกๅชๅˆๅนถไธ€ๅฏนใ€‚

See BPE clearly with an example

e.g. “a b c a b c c”

 Units: a, b, c, a, b, c, c

Step 1: Count all adjacent pairs

Adjacent PairEN๏ผšFrequency
(a, b)2
(b, c)2
(c, a)1
(c, c)1

Highest frequency is (a, b) and (b, c), both 2 times. Pick (a, b) to merge.

(a, b) โ†’ ab

Result๏ผšab c ab c c

Step 2: Count adjacent pairs again๏ผš

PairCount
(ab, c)2
(c, ab)1
(c, c)1

 Highest frequency is (ab, c) with 2 occurrences. Merge.

Result: abc abc c

Step 3 (optional)

Count adjacent pairs again: does (abc, abc) appear?

Check: abc abc c โ†’ adjacent pairs:

  • (abc, abc): 1 occurrence
  • (abc, c): 1 occurrence

(abc, abc) merge into abcabc

Final result comparison

CN๏ผšๆญฅ้ชคEN๏ผšStepCN๏ผš็ป“ๆžœEN๏ผšResult
ๅผ€ๅง‹Starta b c a b c c (7 ไธชๅ•ไฝ)a b c a b c c (7 units)
็ฌฌ 1 ๆญฅๅŽAfter step 1ab c ab c c (5 ไธชๅ•ไฝ)ab c ab c c (5 units)
็ฌฌ 2 ๆญฅๅŽAfter step 2abc abc c (3 ไธชๅ•ไฝ)abc abc c (3 units)

Core summary

CNEN
ๆฏไธ€ๆญฅๅชๅˆๅนถ็›ธ้‚ป็š„ไธคไธชๅ•ไฝEach step merges only two adjacent units
ๅˆๅนถๅŽๅ•ไฝๅ˜ๅฐ‘Units decrease after each merge
ๆ–ฐๅ•ไฝๅฏไปฅๅ‚ไธŽไธ‹ไธ€ๆญฅ็š„ๅˆๅนถNew units can participate in next step’s merges
ๅœๆญขๆกไปถ๏ผš่พพๅˆฐ็›ฎๆ ‡่ฏ่กจๅคงๅฐStop condition: target vocabulary size reached

What is Completion in AI/LLM?

Completion is the fundamental, raw operation of an LLM where the model takes an input text prompt and generates the most likely continuation of that text, token by token, in an autoregressive manner. It has no concept of roles or conversation history โ€” just text in, text out.

Completion ๆ˜ฏ LLM ๆœ€ๅŸบ็ก€ใ€ๆœ€ๅŽŸๅง‹็š„ๆ“ไฝœ๏ผšๆจกๅž‹ๆŽฅๆ”ถไธ€ๆฎต่พ“ๅ…ฅๆ–‡ๆœฌๆ็คบ๏ผŒ็„ถๅŽไปฅ่‡ชๅ›žๅฝ’็š„ๆ–นๅผ้€ token ็”Ÿๆˆ่ฏฅๆ–‡ๆœฌๆœ€ๅฏ่ƒฝ็š„ๅปถ็ปญๅ†…ๅฎนใ€‚ๅฎƒๆฒกๆœ‰่ง’่‰ฒๆˆ–ๅฏน่ฏๅކๅฒ็š„ๆฆ‚ๅฟต โ€” ไป…ไป…ๆ˜ฏๆ–‡ๆœฌ่พ“ๅ…ฅใ€ๆ–‡ๆœฌ่พ“ๅ‡บใ€‚

Key Characteristics (ๅ…ณ้”ฎ็‰นๅพ)

AspectEnglishChinese
InputSingle string promptๅ•ไธชๅญ—็ฌฆไธฒๆ็คบ่ฏ
OutputRaw text continuationๅŽŸๅง‹ๆ–‡ๆœฌๅปถ็ปญ
RolesNoneๆ— 
HistoryMust be manually managedๅฟ…้กปๆ‰‹ๅŠจ็ฎก็†
Underlying mechanismAutoregressive token prediction่‡ชๅ›žๅฝ’ token ้ข„ๆต‹
Modern statusLegacy (GPT-3, Davinci era)้—็•™ๆจกๅผ๏ผˆGPT-3ใ€Davinci ๆ—ถไปฃ๏ผ‰

Simple Example

Prompt (ๆ็คบ่ฏ):     "The capital of France is"
Completion (่กฅๅ…จ):   " Paris."

Prompt (ๆ็คบ่ฏ):     "def fibonacci(n):"
Completion (่กฅๅ…จ):   "\n    if n <= 1:\n        return n\n    else:\n        return fibonacci(n-1) + fibonacci(n-2)"

What is Chat in AI/LLM?

Chat is a structured, turn-based interaction paradigm built on top of completion. It adds role awareness (system, user, assistant) and automatic conversation history management. Each chat interaction is internally converted into a completion with special formatting tokens.

Chatๆ˜ฏๆž„ๅปบๅœจCompletionไน‹ไธŠ็š„็ป“ๆž„ๅŒ–ใ€ๅŸบไบŽ่ฝฎๆฌก็š„ไบคไบ’่Œƒๅผใ€‚ๅฎƒๅขžๅŠ ไบ†่ง’่‰ฒๆ„Ÿ็Ÿฅ๏ผˆ็ณป็ปŸใ€็”จๆˆทใ€ๅŠฉๆ‰‹๏ผ‰ๅ’Œ่‡ชๅŠจๅฏน่ฏๅކๅฒ็ฎก็†ใ€‚ๆฏๆฌกๅฏน่ฏไบคไบ’ๅœจๅ†…้ƒจ้ƒฝ่ขซ่ฝฌๆขไธบๅธฆๆœ‰็‰นๆฎŠๆ ผๅผๆ ‡่ฎฐ็š„่กฅๅ…จใ€‚

Key Characteristics

AspectEnglishChinese
InputArray of messages with rolesๅธฆ่ง’่‰ฒ็š„ๆถˆๆฏๆ•ฐ็ป„
OutputRole-labeled assistant responseๅธฆ่ง’่‰ฒๆ ‡็ญพ็š„ๅŠฉๆ‰‹ๅ›žๅค
RolesSystem, User, Assistant็ณป็ปŸใ€็”จๆˆทใ€ๅŠฉๆ‰‹
HistoryAutomatically managed in message arrayๅœจๆถˆๆฏๆ•ฐ็ป„ไธญ่‡ชๅŠจ็ฎก็†
Underlying mechanismStill completion (with special tokens)ไป็„ถๆ˜ฏ่กฅๅ…จ๏ผˆๅธฆ็‰นๆฎŠๆ ‡่ฎฐ๏ผ‰
Modern statusStandard (GPT-4, Claude, DeepSeek)ๆ ‡ๅ‡†ๆจกๅผ๏ผˆGPT-4ใ€Claudeใ€DeepSeek๏ผ‰
e.g.
messages = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "What is the capital of France?"},
    {"role": "assistant", "content": "The capital of France is Paris."}
]

Completion vs. Chat

Context – understanding the word in Chinese

Context ็š„ๆ ธๅฟƒๆ„ๆ€ๅ…ถๅฎžๆ˜ฏ๏ผšๆจกๅž‹ๅœจ็”Ÿๆˆๅ›ž็ญ”ๆ—ถๆ‰€ไพๆฎ็š„ใ€ๅฏน่ฏๆˆ–ไปปๅŠกไธญๅทฒ็ปๅญ˜ๅœจ็š„ๅ…จ้ƒจๆœ‰ๆ•ˆไฟกๆฏ๏ผˆๅŒ…ๆ‹ฌๅކๅฒๅฏน่ฏใ€ๅฝ“ๅ‰้—ฎ้ข˜ใ€้šๅซๆกไปถใ€็”จๆˆทๅๅฅฝ็ญ‰๏ผ‰ใ€‚่ฟ™ไธช่ฏไฝœไธบ โ€œ่ฏญๅขƒโ€๏ผˆๆœ€ๆŽจ่๏ผ‰,โ€œ่ƒŒๆ™ฏไฟกๆฏโ€ , โ€œๅ‰ๆ–‡่ƒŒๆ™ฏโ€๏ผŒ โ€œๅ…ณ่”ไฟกๆฏโ€ , โ€œไพๆ‰˜ไฟกๆฏโ€, โ€œๅฏน่ฏ่ฎฐๅฟ†โ€๏ผˆ้’ˆๅฏนๅฏน่ฏ็ณป็ปŸ๏ผ‰ ๆฏ”่พƒๅฅฝๅฏนๅบ”ไธญๆ–‡ใ€‚

ๆˆ‘ไธชไบบ่ง‰ๅพ—โ€œ่ฏญๅขƒโ€ๆฏ”่พƒๅฅฝใ€‚

What is Context window๏ผŸ

A Context Window is the amount of information an LLM can โ€œseeโ€ or โ€œrememberโ€ during a conversation or request. Think of it as: The AI model’s working memory. Everything inside the context window can influence the AIโ€™s response.
ๆจกๅž‹ๅช่ƒฝๅŸบไบŽโ€œContext Window ๅ†…็š„ไฟกๆฏโ€ๆฅๅ›ž็ญ”้—ฎ้ข˜๏ผŒ LLM ไธ€ๆฌก่ƒฝโ€œ็œ‹ๅˆฐ/่ฎฐไฝโ€็š„ไฟกๆฏ้‡ๅฐฑๆ˜ฏContext windowsใ€‚ๅฏไปฅ็†่งฃไธบAI็š„โ€ๅ†…ๅญ˜ๅฎน้‡โ€œ

Important Understanding

The context window includes BOTH:

Included in Context WindowExamples
Input tokensprompts, chat history, RAG docs
Output tokensmodel response / AI ่พ“ๅ‡บ๏ผŒๅ›ž็ญ”
Total tokens = Input + Output

Context Engineering

Meaning:

  • deciding WHAT information goes into the context window
  • optimizing token usage
  • ranking retrieved documents
  • summarizing history
  • removing irrelevant content

Context Engineering ไนŸๅฐฑๆ˜ฏ๏ผšโ€œๅ†ณๅฎšไป€ไนˆไฟกๆฏ่ฟ›ๅ…ฅ Context Windowโ€ใ€‚ๅŒ…ๆ‹ฌ๏ผš

  • ๅ“ชไบ›ๆ–‡ๆกฃๆœ€้‡่ฆ
  • ๅฆ‚ไฝ•่Š‚็œ token
  • ๅฆ‚ไฝ•ๅŽ‹็ผฉๅކๅฒ
  • ๅฆ‚ไฝ•ๆŽ’ๅบ RAG ็ป“ๆžœ
  • ๅฆ‚ไฝ•ๅŽปๆމๆ— ๅ…ณไฟกๆฏ

่ฟ™ๆ˜ฏ Enterprise AI ้žๅธธๆ ธๅฟƒ็š„่ƒฝๅŠ›ใ€‚


Deployment = making the model callable/useable.

Without deployment:

  • model exists in catalog
  • but your app cannot use it

After deployment, Azure gives endpoint + API access

VERY important

Deployment โ‰  Agent

A deployment is: an exposed model service
ๅฏไปฅ็ฎ€ๅ•็†่งฃไธบ๏ผšๅฐ†ไธ€ไธชModel, ไพ‹ๅฆ‚ๅฐ† GPT ๆˆ– Deep Seek๏ผŒ่ฐƒ่ฟ›ๆˆ‘็š„็ณป็ปŸ๏ผŒๅนถๆฟ€ๆดปๅฎƒ๏ผŒ่ฎฉ่ฟ™ไธชmodel ๅœจๆˆ‘็š„็ณป็ปŸ้‡Œๅ˜ไธบโ€œๅฏไฝฟ็”จไบ†โ€ใ€‚

What are Embeddings in AI / LLM?

Embeddings in AI / LLM are numerical representations of text (or other data like images, audio) in a highโ€‘dimensional vector space. Simply put, they turn words, sentences, or documents into lists of numbers so that computers can โ€œunderstandโ€ their meaning mathematically.

ๅœจ AI / ๅคง่ฏญ่จ€ๆจกๅž‹ไธญ๏ผŒEmbedding ๆ˜ฏๆŠŠๆ–‡ๆœฌ๏ผˆๆˆ–ๅ›พๅƒใ€้Ÿณ้ข‘็ญ‰๏ผ‰่ฝฌๆขๆˆๆ•ฐๅญ—ๅˆ—่กจ๏ผˆๅ‘้‡๏ผ‰ ็š„ๆŠ€ๆœฏใ€‚็ฎ€ๅ•่ฏด๏ผŒๅฐฑๆ˜ฏ่ฎฉ่ฎก็ฎ—ๆœบ้€š่ฟ‡ไธ€ไธฒๆ•ฐๅญ—ๆฅโ€œ็†่งฃโ€ๆ–‡ๅญ—็š„ๅซไน‰ใ€‚

Key points:

  • What it looks like:
    A word like "king" might be represented as a vector:
    [0.25, -0.78, 0.43, โ€ฆ, 0.12] (e.g., 300โ€“4096 dimensions).
  • How it works:
    Words or phrases with similar meanings are placed close together in this vector space.
    • "king" and "queen" are close.
    • "apple" (fruit) and "apple" (company) have different vectors depending on context.
  • Why embeddings matter:
    • They capture semantic meaning โ€“ relationships like king โˆ’ man + woman โ‰ˆ queen.
    • They enable search (find similar texts), clustering (group topics), and recommendation.
    • LLMs use embeddings internally to process every token you feed into the model.

Vector Databases

A Vector Database is a database designed to store and search embeddings (vectors). Vector DB stores semantic meaning vectors

Common Vector Databases

  • Pinecone / ๅ…จๆ‰˜็ฎกใ€ๆ— ๆœๅŠกๅ™จใ€ไฝŽๅปถ่ฟŸ
  • Weaviate / ๅ†…็ฝฎๆททๅˆๆœ็ดข + ๆจกๅ—ๅŒ–
  • FAISS / ๅบ“๏ผˆ้žๆ•ฐๆฎๅบ“๏ผ‰๏ผŒ้ซ˜ๅบฆไผ˜ๅŒ–็š„ANN
  • Azure AI Search
  • Databricks Vector Search
  • Milvus / ไบ‘ๅŽŸ็”Ÿใ€GPUๅŠ ้€Ÿใ€ๅไบฟ็บง่ง„ๆจก
  • Chroma / ่ฝป้‡็บงใ€ๅตŒๅ…ฅๅผใ€ๅŽŸ็”ŸPython

These databases optimize:
nearest neighbor search
semantic retrieval
high-dimensional vector operations

Traditional Database vs Vector Database

Traditional DatabaseVector Database
Stores rows/columnsStores vectors
SQL queriesSimilarity search
Exact matchingSemantic matching
Keyword searchMeaning search
Structured dataEmbeddings

Example, Suppose company documents contain: “Employees may work remotely twice weekly.”

User asks: “What is the work from home policy?”. Traditional keyword search may fail because โ€œremoteโ€ โ‰  โ€œwork from homeโ€. But embedding vectors capture semantic similarity.

Semantic Search

Similarity search is a technique that finds items in a dataset that are most similar to a given query vector, based on distance metrics in a high-dimensional embedding space โ€” enabling semantic matching rather than exact keyword matching.
็›ธไผผๆ€งๆœ็ดข(Similarity search) ๆ˜ฏไธ€็งๆŠ€ๆœฏ๏ผŒๅŸบไบŽ้ซ˜็ปดๅตŒๅ…ฅ็ฉบ้—ดไธญ็š„่ท็ฆปๅบฆ้‡๏ผŒๅœจๆ•ฐๆฎ้›†ไธญๆ‰พๅˆฐไธŽ็ป™ๅฎšๆŸฅ่ฏขๅ‘้‡ๆœ€็›ธไผผ็š„้กน็›ฎ โ€” ๅฎž็Žฐ่ฏญไน‰ๅŒน้…่€Œ้ž็ฒพ็กฎๅ…ณ้”ฎ่ฏๅŒน้…ใ€‚


Grounding = making the AI answer based on real external evidence, not memory.
่ฎฉ AI ็š„ๅ›ž็ญ”โ€œๆœ‰ไพๆฎโ€๏ผŒไธๆ˜ฏ้ ่ฎฐๅฟ†ไนฑ็Œœใ€‚ๆˆ–่€…่ฏดโ€œ็ป™ AI ็œ‹่ต„ๆ–™โ€๏ผŒ ไธๆ˜ฏ่ฎฉๅฎƒ่‡ชๅทฑๆƒณ็ญ”ๆกˆ

What is Hallucination in AI / LLM?

Hallucination in AI / LLM refers to the phenomenon where the model generates content that is factually incorrect, nonsensical, or completely unrelated to the real world or the provided source, while presenting it with high confidence as if it were true. This is one of the BIGGEST concerns in enterprise AI systems.

Common examples include:

  • Inventing nonโ€‘existent references, laws, or historical events.
  • Incorrectly calculating simple arithmetic.
  • Misinterpreting the userโ€™s input and fabricating plausibleโ€‘sounding but false information.

่™šๅ‡็”Ÿๆˆ, ๆจกๅž‹็ผ–้€  AI ็”Ÿๆˆไบ†้”™่ฏฏ็š„๏ผŒ็ผ–้€ ็š„๏ผŒไธ็œŸๅฎž็š„๏ผŒๆฒกไพๆฎ็š„็š„ไฟกๆฏใ€‚ไฝ†AIๅดโ€œๅพˆ่‡ชไฟกโ€ๅœฐ่ฏดๅ‡บๆฅใ€‚ๅณ๏ผš ๆจกๅž‹่‡ชไฟกๅœฐ่พ“ๅ‡บ้”™่ฏฏๆˆ–ๅ‡ญ็ฉบๆ้€ ็š„ไฟกๆฏ

Why Hallucinations Happen๏ผŸ

LLMs Predict Language, Not Truth๏ผ›
2) Missing Context. If the model lacks:
  • sufficient information
  • enterprise data
  • current data

it may โ€œfill in the gaps.โ€

3) Ambiguous Prompts. Poor prompts can cause:
  • assumptions
  • invented details
  • unstable outputs
4) Outdated Training Data. Models have training cutoffs. They may:
  • not know recent events
  • generate outdated answers
  • guess newer information
5) Weak RAG / Retrieval. In enterprise AI:
  • bad retrieval
  • irrelevant documents
  • incomplete grounding

can produce hallucinated answers.

Types of Hallucinations

  • A. Factual Hallucination: Wrong facts. /ไบ‹ๅฎžๅนป่ง‰, ไบ‹ๅฎž้”™่ฏฏ, ็ผ–้€ ๅ…ฌๅธๆ”ฟ็ญ–
  • B. Citation Hallucination: Fake sources or references. / ๅผ•็”จๅนป่ง‰, ๅ‡่ฎบๆ–‡ใ€ๅ‡ๆฅๆบใ€‚
  • C. Logical Hallucination: Reasoning errors / ๆŽจ็†ๅนป่ง‰, ้€ป่พ‘ๆŽจ็†้”™่ฏฏ
  • D. Tool/API Hallucination: Inventing APIs, functions, parameters, libraries / ็ผ–้€ API็ญ‰

How Enterprises Reduce Hallucinations

1) RAG (Retrieval-Augmented Generation)
  • RAG: Most important technique. Instead of relying only on model memory / ๆœ€ๆ ธๅฟƒ
  • Better Prompt Engineering, Clear prompts reduce ambiguity.
  • Context Engineering Control: what information enters context, retrieval quality, ranking, chunking, summarization.
  • Evaluation Systems: AI outputs are tested for: factual accuracy, roundedness, consistency, safety.
  • Human-in-the-Loop, Humans validate :sensitive outputs, approvals, critical decisions.


What is LLM?
Large language models, also known as LLMs, are very large deep learning models that are pre-trained on vast amounts of data. The underlying transformer is a set of neural networks that consist of an encoder and a decoder with self-attention capabilities. The encoder and decoder extract meanings from a sequence of text and understand the relationships between words and phrases in it.

ๅคงๅž‹่ฏญ่จ€ๆจกๅž‹๏ผˆ่‹ฑ่ฏญ๏ผšlarge language model๏ผŒLLM๏ผ‰๏ผŒไนŸ็งฐๅคง่ฏญ่จ€ๆจกๅž‹๏ผŒ็ฎ€็งฐๅคงๆจกๅž‹๏ผŒๆ˜ฏไธ€็งๅŸบไบŽไบบๅทฅ็ฅž็ป็ฝ‘็ปœ็š„ๅทฒ็ป่ฎญ็ปƒ่ฟ‡็š„่ฏญ่จ€ๆจกๅž‹ใ€‚ๅคง่ฏญ่จ€ๆจกๅž‹ไธ“ไธบ่‡ช็„ถ่ฏญ่จ€ๅค„็†ไปปๅŠก่€Œ่ฎพ่ฎก๏ผŒๅฐคๅ…ถ้€‚็”จไบŽ่ฏญ่จ€็”Ÿๆˆใ€‚

ไป–ไปฌๅ…ณ็ณปๅŸบๆœฌๅฆ‚่ฟ™ไธชๅฑ‚็บง็ป“ๆž„/ๅŒ…ๅซๅ…ณ็ณป๏ผš
ไบบๅทฅๆ™บ่ƒฝ (AI) > ๆจกๅž‹ (Model) > ็”Ÿๆˆๅผ AI (Generative AI) > ๅคง่ฏญ่จ€ๆจกๅž‹ (LLM)


Memory = system that stores user/context over time

Types:

๐Ÿ”น Short-term memory
  • current conversation context
๐Ÿ”น Long-term memory
  • user preferences
  • past interactions
  • profile data

Why important?
  • every chat is โ€œresetโ€ / ๆฒกๆœ‰memory,ๆฏๆฌก้ƒฝๆ˜ฏๆ–ฐ็”จๆˆท
  • personalized AI experience / ๆœ‰memory AI ๅ˜ๆˆโ€œไธชไบบๅŠฉ็†โ€

What is Prompt? A Prompt is the instruction, question, context, or input you give to an AI model (LLM) to tell it what you want it to do.
ๅฐฑๆ˜ฏไฝ ็ป™ AI ็š„โ€œๆŒ‡ไปค/่พ“ๅ…ฅโ€๏ผŒ ๅ‘Š่ฏ‰AI๏ผš ่ฆๅšไป€ไนˆ๏ผŒ็”จไป€ไนˆๆ–นๆณ•ๅš๏ผŒ่พ“ๅ‡บไป€ไนˆใ€‚

e.g.

Summarize this document in 5 bullet points.

That sentence is a prompt.

Another example:

You are a senior Azure architect.
Explain Medallion Architecture for a banking platform.

The prompt tells the AI:

  • its role
  • the task
  • the expected output
  • sometimes the tone/style

Basic Prompt Structure

A prompt often contains:

PartPurpose
InstructionWhat to do / ๅšไป€ไนˆ
ContextBackground information / ่ƒŒๆ™ฏไฟกๆฏ
ConstraintsRules/limits/ ้™ๅˆถๆกไปถ
ExamplesDemonstrations / ็คบไพ‹
Output formatExpected response structure / ่ฆๆฑ‚็š„่พ“ๅ‡บๆ ผๅผ

Example

You are a data architect.

Context:
The company uses Azure Databricks and Synapse.

Task:
Design a metadata-driven ingestion framework.

Output:
Provide architecture, components, and best practices.

This is a more structured prompt.

Core components of Good prompt

A good prompt usually includs:

ENCN
Goal็›ฎๆ ‡
Context่ƒŒๆ™ฏ
Constraints้™ๅˆถๆกไปถ
Input่พ“ๅ…ฅๆ•ฐๆฎ
Output Format่พ“ๅ‡บๆ ผๅผ
Examples็คบไพ‹

What is Prompt Engineering?

Prompt Engineering = the practice of designing prompts to get better AI outputs.
ๆ็คบ่ฏๅทฅ็จ‹ๅฐฑๆ˜ฏ๏ผš่ฎพ่ฎก Prompt ๆฅ่Žทๅพ—ๆ›ดๅฅฝ AI ่พ“ๅ‡บโ€็š„ๆŠ€ๆœฏใ€‚
ๅ…ถๅฎž็ฎ€ๅ•่ฏดๅฐฑๆ˜ฏ โ€œไผš้—ฎ AI ้—ฎ้ข˜โ€

It is:

  • writing prompts strategically / ๆ›ด่ชๆ˜Žๅœฐๅ†™ Prompt
  • structuring context correctly / ๆ›ดๅˆ็†ๅœฐ็ป„็ป‡ Context
  • controlling AI behavior / ๆ›ด็จณๅฎšๅœฐๆŽงๅˆถ AI ่กŒไธบ
  • improving reliability and quality / ๆ้ซ˜ๅฏ้ ๆ€งๅ’Œ่ดจ้‡

Think of it as:

Programming with language instead of code.
ๅฏ็†่งฃๆˆ๏ผš็”จ่‡ช็„ถ่ฏญ่จ€โ€œ็ผ–็จ‹โ€๏ผŒ่€Œไธๆ˜ฏ็”จไปฃ็ ็ผ–็จ‹

Why Prompt Engineering Matters / ไธบไป€ไนˆ้‡่ฆ

LLMs are highly sensitive to / LLMๅฏน่ฟ™ไบ›้ซ˜ๅบฆๆ•ๆ„Ÿ:

  • wording / ๆŽช่พž
  • context / ่ƒŒๆ™ฏไฟกๆฏ
  • instructions / ่ฆๆฑ‚
  • examples / ็คบไพ‹
  • formatting / ่พ“ๅ‡บ่ฆๆฑ‚

Small prompt changes can dramatically affect / ๅฏน promptไธŠ่ฟฐ่ฟ™ไบ›ๅ“ชๆ€•ๆ˜ฏๅฐ็š„ๆ”นๅŠจ้ƒฝไผšๅฝฑๅ“ๅˆฐ็ป“ๆžœ:

  • accuracy / ๅ‡†็กฎ็އ
  • reasoning / ๆŽจ็†่ƒฝๅŠ›
  • hallucination / ๅนป่ง‰, ๆ— ๆ นๆฎ็š„็ป“่ฎบ
  • consistency / ็จณๅฎšๆ€ง
  • output quality / ่พ“ๅ‡บ่ดจ้‡

Common Prompt Engineering Techniques

common prompt technical include:

TechCN
Role PromptingๆŒ‡ๅฎš AI ่บซไปฝ
Few-shot็ป™ๅคšไธชไพ‹ๅญ
Chain of Thoughtๅผ•ๅฏผ AI ไธ€ๆญฅไธ€ๆญฅๆ€่€ƒ
Output ControlๆŽงๅˆถ่พ“ๅ‡บๆ ผๅผ
ConstraintsๅŠ ้™ๅˆถๆกไปถ
Context Injectionๆณจๅ…ฅไธšๅŠก่ƒŒๆ™ฏ

1) Role Prompting

Tell the AI who it is.

Example:

You are a senior enterprise architect.

This changes response style and depth.


2) Context Injection

Provide necessary information / ๆไพ›/ๆณจๅ…ฅๅฟ…่ฆ็š„่ƒŒๆ™ฏไฟกๆฏ๏ผŒไปฅๆ้ซ˜็ป“ๆžœ็š„ๅ‡†็กฎๆ€ง

Example:

The environment uses:
- Azure Databricks
- Delta Lake
- Unity Catalog

Without context, AI guesses / ไธๆไพ›่ฟ™ไบ›่ƒŒๆ™ฏ่ต„ๆ–™๏ผŒAIไผšๅŽปไนฑ็Œœใ€‚ๅฝฑๅ“็ป“ๆžœ็š„ๅ‡†็กฎๆ€ง.


3) Output Formatting

Specify desired structure / ่พ“ๅ‡บๆ ผๅผๆŽงๅˆถ๏ผŒ ็ป™AIๆๅ‡บ่พ“ๅ‡บ็š„ๆ ผๅผ่ฆๆฑ‚๏ผŒ ๅฏไปฅๅธฎๅŠฉๆ้ซ˜็ป“ๆžœ็š„ๅ‡†็กฎๆ€ง

Example:

Return the answer as:
- architecture diagram
- bullet points
- implementation steps

4) Few-Shot Prompting

Give examples of desired behavior.
Few-Shot Learning / (ๅฐ‘ๆ ทๆœฌๅญฆไน ) ๆ˜ฏไธ€็งไบบๅทฅๆ™บ่ƒฝๆŠ€ๆœฏใ€‚ๅฎƒๆŒ‡็š„ๆ˜ฏๅœจ็ป™ๆจกๅž‹็š„ๆ็คบ่ฏ๏ผˆPrompt๏ผ‰ไธญๆไพ›ๅฐ‘้‡๏ผˆ้€šๅธธ 2 ๅˆฐ 5 ไธช๏ผ‰็คบไพ‹๏ผŒๅธฎๅŠฉๆจกๅž‹็†่งฃไปปๅŠก่ฆๆฑ‚๏ผŒไปŽ่€Œ็”Ÿๆˆๆ›ดๅ‡†็กฎ็š„ๅ›žๅคใ€‚

Example:

I want to classify sentiment.
Example 1: "I love this food!" -> Positive
Example 2: "This is the worst day ever." -> Negative
Example 3: "The movie was okay." -> Neutral

Now classify this: "The weather is quite nice today." ->

output:  Positive

AI learns pattern/style from examples.


5) Chain-of-Thought Prompting

Chain-of-Thought (CoT) is a prompting technique that forces the LLM to show its reasoning steps before giving a final answer โ€” like asking a SQL analyst to explain their logic before writing the query.

It transforms input โ†’ answer into input โ†’ step1 โ†’ step2 โ†’ โ€ฆ โ†’ answer.

CoT ๆ˜ฏไธ€็งๆ็คบๆŠ€ๆœฏ๏ผŒๅผบ่ฟซ LLM ๅœจ็ป™ๅ‡บๆœ€็ปˆ็ญ”ๆกˆไน‹ๅ‰ๅฑ•็คบๆŽจ็†ๆญฅ้ชค โ€”โ€” ๅฐฑๅƒ่ฎฉๆ•ฐๆฎๅˆ†ๆžๅธˆๅœจๅ†™ SQL ไน‹ๅ‰ๅ…ˆ่งฃ้‡Š้€ป่พ‘ไธ€ๆ ทใ€‚
ๅฎƒๅฐ† ่พ“ๅ…ฅ โ†’ ็ญ”ๆกˆ  ่ฝฌๅ˜ไธบ ่พ“ๅ…ฅ โ†’ ๆญฅ้ชค1 โ†’ ๆญฅ้ชค2 โ†’ โ€ฆ โ†’ ็ญ”ๆกˆ

Example:

# Question๏ผš 
"Roger has 5 marbles. He buys 2 bags with 4 marbles each. Then he loses 3 marbles. Think step by step. How many does he have left?"



# the answer with all steps: 

Roger starts with 5 marbles.
He buys 2 bags with 4 marbles each:

2 ร— 4 = 8 marbles

Now he has:

5 + 8 = 13 marbles

Then he loses 3 marbles:

13 โˆ’ 3 = 10 marbles

Answer: 10 marbles. ๐ŸŸข

e.g.

# Question
ไธ€ไธชๅ†œๅคซๆœ‰15ๅช้ธกใ€‚ไธ€ๅช็‹็‹ธๆฏๆ™šๅƒๆމ3ๅช้ธก๏ผŒ่ฟž็ปญ2ๆ™šใ€‚็„ถๅŽๅ†œๅคซๅˆไนฐไบ†5ๅช้ธกใ€‚็ฌฌไบŒๅคฉๆ™šไธŠ๏ผŒ็‹็‹ธๅƒๆމ2ๅช้ธกใ€‚่ฟ˜ๅ‰ฉๅคšๅฐ‘ๅช้ธก๏ผŸ


# result 
ๆˆ‘ไปฌไธ€ๆญฅไธ€ๆญฅ็ฎ—๏ผš

ๅˆๅง‹ๆ•ฐ้‡
ๅ†œๅคซๆœ‰ 15 ๅช้ธก

็ฌฌ1ๆ™š็‹็‹ธๅƒๆމ 3 ๅช
15 โˆ’ 3 = 12

็ฌฌ2ๆ™š็‹็‹ธๅ†ๅƒๆމ 3 ๅช
12 โˆ’ 3 = 9

ๅ†œๅคซๅˆไนฐไบ† 5 ๅช้ธก
9 + 5 = 14

็ฌฌไบŒๅคฉๆ™šไธŠ็‹็‹ธๅˆๅƒๆމ 2 ๅช
14 โˆ’ 2 = 12



Final Answer:
12 chickens remain

Useful for:

  • logic
  • architecture
  • math
  • troubleshooting

Prompt Injection

Prompt Injection is an attack where malicious user input tries to override or hijack the system prompt, making the AI behave in unintended ways. In production, prompt injection is one of the Top 5 LLM security risks (OWASP LLM Top 10). Any customer-facing AI must implement these defenses.

ๆ็คบ่ฏๆณจๅ…ฅๆ˜ฏไธ€็งๆ”ปๅ‡ปๆ–นๅผ๏ผŒๆถๆ„็”จๆˆท่พ“ๅ…ฅ่ฏ•ๅ›พ่ฆ†็›–ๆˆ–ๅŠซๆŒ system prompt๏ผŒ่ฎฉAIๅšๅ‡บ้ž้ข„ๆœŸ็š„่กŒไธบใ€‚


Schema enforcement = forcing data to follow a fixed structure (schema), not free-form text.

ๅผบๅˆถ AI ๆˆ– API ่พ“ๅ‡บโ€œ็ฌฆๅˆๆ ผๅผ็š„ๆ•ฐๆฎโ€๏ผŒไธ่ƒฝไนฑๅ†™ใ€‚ไพ‹ๅฆ‚
John is 30 years old and lives in Toronto

AI ๅฏ่ƒฝ่พ“ๅ‡บ๏ผšJohn is 30 years old and lives in Toronto
ไนŸๅฏ่ƒฝๆ˜ฏ๏ผš
name: John
age: thirty
location: Toronto Canada maybe

ไธ็จณๅฎšใ€ไธๅฏๆœบๅ™จๅค„็†ใ€‚ ็”จSchema enforcement ๅผบๅˆถๅฎƒ่พ“ๅ‡บ่ฟ™ๆ ท
{
“name”: “John”,
“age”: 30,
“location”: “Toronto”
}

StreamingIn the context of Large Language Models (LLMs), streaming refers to the technique of returning generated tokens one by one (or in small chunks) as soon as they are produced by the model, rather than waiting for the entire response to be completed. The underlying transport is typically Server-Sent Events (SSE) or chunked HTTP responses, where the server pushes incremental updates to the client.

Streaming

async/await LLM Call

What is Temperature?

Temperature is a hyperparameter that controls the randomness or creativity of an LLM’s output. It scales the logits (raw prediction scores) before the softmax function that converts them into probabilities โ€” lower temperatures make the model more deterministic and focused, while higher temperatures make it more diverse and exploratory.
Temperature ๆ˜ฏไธ€ไธช่ถ…ๅ‚ๆ•ฐ๏ผŒ็”จไบŽๆŽงๅˆถๅคง่ฏญ่จ€ๆจกๅž‹่พ“ๅ‡บ็š„้šๆœบๆ€งๆˆ–ๅˆ›้€ ๆ€งใ€‚ๅฎƒๅœจ softmax ๅ‡ฝๆ•ฐ๏ผˆๅฐ†ๅŽŸๅง‹้ข„ๆต‹ๅˆ†ๆ•ฐ่ฝฌๆขไธบๆฆ‚็އ๏ผ‰ไน‹ๅ‰ๅฏน่ฟ™ไบ› logits ่ฟ›่กŒ็ผฉๆ”พ โ€” ่พƒไฝŽ็š„ๆธฉๅบฆไฝฟๆจกๅž‹ๆ›ด็กฎๅฎšใ€ๆ›ดไธ“ๆณจ๏ผŒ่€Œ่พƒ้ซ˜็š„ๆธฉๅบฆไฝฟๅ…ถๆ›ดๅคšๆ ทๅŒ–ใ€ๆ›ดๅ…ทๆŽข็ดขๆ€งใ€‚

 Imagine you’re at a restaurant with a menu of 10 dishes. Temperature controls how likely you are to pick your absolute favorite vs. trying something new.
 ๆƒณ่ฑกไฝ ๅœจไธ€ไธชๆœ‰ 10 ้“่œ็š„้คๅŽ…้‡Œใ€‚ๆธฉๅบฆๆŽงๅˆถ็€ไฝ ้€‰ๆ‹ฉๆœ€็ˆฑ็š„่œ vs. ๅฐ่ฏ•ๆ–ฐ่œ็š„ๅฏ่ƒฝๆ€งใ€‚

TemperatureAnalogy (English)Analogy (ไธญๆ–‡)
Low (0.1 ~ 0.3)You always order your #1 favorite dish. Very predictable.ไฝ ๆ€ปๆ˜ฏ็‚นไฝ ๆœ€็ˆฑ็š„็ฌฌไธ€้“่œใ€‚้žๅธธๅฏ้ข„ๆต‹ใ€‚
Medium (0.7 ~ 1.0)You usually pick your top dish, but sometimes try #2 or #3. Balanced.ไฝ ้€šๅธธ้€‰ๆœ€็ˆฑ็š„่œ๏ผŒไฝ†ๆœ‰ๆ—ถๅฐ่ฏ•็ฌฌไบŒๆˆ–็ฌฌไธ‰ๅ–œๆฌข็š„ใ€‚ๅนณ่กกใ€‚
High (1.5+)You randomly pick any dish, even ones you don’t know. Very unpredictable.ไฝ ้šๆœบ้€‰ไปปไฝ•่œ๏ผŒ็”š่‡ณไฝ ไธ่ฎค่ฏ†็š„่œใ€‚้žๅธธไธๅฏ้ข„ๆต‹ใ€‚

Top-K โ€“ Sample only from the K most probable tokens

  • The model looks at all possible next tokens and their probabilities.
  • It keeps only the K tokens with the highest probabilities and discards the rest.
  • Then it randomly selects one token from these K tokens (using their relative probabilities).

Effect:

  • Smaller K (e.g., 10) โ†’ Fewer choices โ†’ More deterministic, predictable, and safe outputs.
  • Larger K (e.g., 100) โ†’ More choices โ†’ More random and diverse outputs.

Example (K=3):
Probabilities: “cat” (50%), “dog” (30%), “bird” (12%), “car” (5%), “tree” (3%)
โ†’ Keep only {cat, dog, bird} โ†’ “car” and “tree” can never be chosen.

ๅชไปŽๆฆ‚็އๆœ€้ซ˜็š„ K ไธช token ไธญ้‡‡ๆ ท

  • ๆจกๅž‹ๅ…ˆ็ฎ—ๅ‡บๆ‰€ๆœ‰ๅฏ่ƒฝ็š„ไธ‹ไธ€ไธช token ๅŠๅ…ถๆฆ‚็އใ€‚
  • ๅชไฟ็•™ ๆฆ‚็އๆœ€้ซ˜็š„ๅ‰ K ไธช token๏ผŒๆ‰”ๆމๅ…ถไฝ™ tokenใ€‚
  • ็„ถๅŽๅœจ่ฟ™ K ไธช token ไธญๆŒ‰ๆฆ‚็އ้šๆœบ้€‰ไธ€ไธชใ€‚

ๆ•ˆๆžœ๏ผš

  • K ่ถŠๅฐ๏ผˆๅฆ‚ 10๏ผ‰ โ†’ ๅฏ้€‰่ฏ่ถŠๅฐ‘ โ†’ ่พ“ๅ‡บ่ถŠ ็กฎๅฎšใ€ๅฎ‰ๅ…จใ€ๅฏ้ข„ๆต‹ใ€‚
  • K ่ถŠๅคง๏ผˆๅฆ‚ 100๏ผ‰ โ†’ ๅฏ้€‰่ฏ่ถŠๅคš โ†’ ่พ“ๅ‡บ่ถŠ ้šๆœบใ€ๅคšๆ ทๅŒ–ใ€‚

ไพ‹ๅญ๏ผˆK=3๏ผ‰๏ผš
ๆฆ‚็އ๏ผš็Œซ(50%)ใ€็‹—(30%)ใ€้ธŸ(12%)ใ€่ฝฆ(5%)ใ€ๆ ‘(3%)
โ†’ ๅชไฟ็•™ {็Œซ, ็‹—, ้ธŸ} โ†’ โ€œ่ฝฆโ€โ€œๆ ‘โ€ ๆฐธ่ฟœไธๅฏ่ƒฝ่ขซ้€‰ไธญใ€‚

Top-P (Nucleus Sampling) โ€“ Choose the smallest set of tokens whose cumulative probability โ‰ฅ P

  • Instead of a fixed number of tokens (K), Top-P dynamically selects tokens from the most probable downward until the sum of their probabilities reaches or exceeds P.
  • This selected set is called the nucleus.

Effect:

  • Smaller P (e.g., 0.9) โ†’ Keeps only the top few high-probability tokens โ†’ More stable.
  • Larger P (e.g., 0.95โ€“1.0) โ†’ Keeps more tokens (sometimes all) โ†’ More random.

Example (P=0.9):
Probabilities: “cat” (50%, cumulative 50%), “dog” (30%, cumulative 80%), “bird” (12%, cumulative 92% โ‰ฅ 90%)
โ†’ Nucleus = {cat, dog, bird} โ†’ “car” and “tree” are excluded.

Key difference from Top-K:
If the probability distribution is very flat, P=0.9 might keep 20+ tokens. If very sharp, it might keep only 1 token. Top-K always keeps exactly K tokens.

Topโ€‘P๏ผˆๆ ธ้‡‡ๆ ท๏ผ‰ โ€“ ้€‰ๆ‹ฉ็ดฏ่ฎกๆฆ‚็އ โ‰ฅ P ็š„ๆœ€ๅฐ token ้›†ๅˆ

  • ไธๅ›บๅฎš token ไธชๆ•ฐ๏ผŒ่€Œๆ˜ฏ ไปŽๆฆ‚็އๆœ€้ซ˜็š„ token ๅผ€ๅง‹ๅพ€ไธ‹ๅŠ ๏ผŒ็›ดๅˆฐ ็ดฏ่ฎกๆฆ‚็އ โ‰ฅ Pใ€‚
  • ่ฟ™ไธชๅŠจๆ€้€‰ๅ‡บๆฅ็š„ token ้›†ๅˆๅซๅš โ€œๆ ธ๏ผˆnucleus๏ผ‰โ€ใ€‚

ๆ•ˆๆžœ๏ผš

  • P ่ถŠๅฐ๏ผˆๅฆ‚ 0.9๏ผ‰ โ†’ ๅชไฟ็•™ๅฐ‘ๆ•ฐๅ‡ ไธช้ซ˜ๆฆ‚็އ token โ†’ ่พ“ๅ‡บ่ถŠ ็จณๅฎšใ€‚
  • P ่ถŠๅคง๏ผˆๅฆ‚ 0.95๏ฝž1.0๏ผ‰ โ†’ ไฟ็•™ๆ›ดๅคš token๏ผˆ็”š่‡ณๅ…จ้ƒจ๏ผ‰ โ†’ ่พ“ๅ‡บ่ถŠ ้šๆœบใ€‚

ไพ‹ๅญ๏ผˆP=0.9๏ผ‰๏ผš
ๆฆ‚็އ๏ผš็Œซ(50%ใ€็ดฏ่ฎก50%)ใ€็‹—(30%ใ€็ดฏ่ฎก80%)ใ€้ธŸ(12%ใ€็ดฏ่ฎก 92% โ‰ฅ 90%)
โ†’ ๅ€™้€‰้›† = {็Œซ, ็‹—, ้ธŸ} โ†’ โ€œ่ฝฆโ€โ€œๆ ‘โ€ ่ขซๆŽ’้™คใ€‚

ไธŽ Top-K ็š„ๅ…ณ้”ฎๅŒบๅˆซ๏ผš
ๅฆ‚ๆžœๆฆ‚็އๅˆ†ๅธƒๅพˆๅนณๅฆ๏ผŒP=0.9 ๅฏ่ƒฝไฟ็•™ 20+ ไธช token๏ผ›ๅฆ‚ๆžœๅพˆๅฐ–้”๏ผŒๅฏ่ƒฝๅชไฟ็•™ 1 ไธช tokenใ€‚
่€Œ Top-K ๆฐธ่ฟœๅ›บๅฎšไฟ็•™ K ไธช tokenใ€‚

Why use them together?

  • Top-K alone: Can still include unlikely tokens if K is large.
  • Top-P alone: Works well alone, but combined with Top-K (e.g., top_k=50, top_p=0.95) โ†’ First limit to K tokens, then apply nucleus โ†’ Best balance.

ไธบไป€ไนˆ่ฆ็ป„ๅˆไฝฟ็”จ๏ผŸ

  • ๅช็”จ Top-K๏ผšK ่พƒๅคงๆ—ถไปๅฏ่ƒฝไฟ็•™ไธๅˆ็†็š„ tokenใ€‚
  • ๅช็”จ Top-P๏ผšๅทฒ็ปๅพˆไธ้”™๏ผŒไฝ†ๅ’Œ Top-K ็ป„ๅˆ๏ผˆๅฆ‚ top_k=50, top_p=0.95๏ผ‰โ†’ ๅ…ˆ้™ๅˆถๆœ€ๅคš 50 ไธช token๏ผŒๅ†ไปŽไธญๆŒ‘ๆ ธ โ†’ ๆ—ขๆŽ’้™คๅฐพ้ƒจๅžƒๅœพ่ฏ๏ผŒๅˆไฟๆŒ็ตๆดปๆ€งใ€‚

What are Tokens in AI / LLM?

Tokens in AI / LLM are the basic units of text that the model reads and generates. Instead of processing raw text characterโ€‘byโ€‘character or wordโ€‘byโ€‘word, the model breaks text into smaller, meaningful pieces called tokens.

ๅœจ AI / ๅคง่ฏญ่จ€ๆจกๅž‹ไธญ๏ผŒToken ๆ˜ฏๆจกๅž‹ๅค„็†ๆ–‡ๆœฌๆ—ถ็š„ๆœ€ๅŸบๆœฌๅ•ๅ…ƒใ€‚ๆจกๅž‹ไธไผšไธ€ไธชๅญ—็ฌฆไธ€ไธชๅญ—็ฌฆๅœฐ่ฏป๏ผŒไนŸไธไผšๆŒ‰ๅฎŒๆ•ดๅ•่ฏ่ฏป๏ผŒ่€Œๆ˜ฏๆŠŠๆ–‡ๆœฌๅˆ‡ๅˆ†ๆˆๆœ‰ๆ„ไน‰็š„็‰‡ๆฎต๏ผŒๆฏไธช็‰‡ๆฎตๅฐฑๆ˜ฏไธ€ไธช Tokenใ€‚

Token ๅฐฑๆ˜ฏๆŠŠไธ€ๅฅ่ฏๅˆ‡ๆˆๆจกๅž‹่ƒฝโ€œๆถˆๅŒ–โ€็š„ๆœ€ๅฐ็ขŽ็‰‡๏ผŒๆฏไธช็ขŽ็‰‡ๆœ‰็›ธๅฏน็‹ฌ็ซ‹็š„ๆ„ไน‰ใ€‚ๅˆ‡็š„ๆ–นๅผๅ–ๅ†ณไบŽๅˆ†่ฏๅ™จ๏ผŒไธๅŒๆจกๅž‹ๅˆ‡ๆณ•ๅฏ่ƒฝไธไธ€ๆ ทใ€‚

Key points:

  • A token is not always a whole word, nor a single character. It can be:
    • A short common word: "cat" โ†’ 1 token
    • Part of a longer word: "unhappiness" โ†’ "un" + "happiness" (2 tokens)
    • A single character: "a" โ†’ 1 token
    • A punctuation mark: "." โ†’ 1 token
    • A space or part of a space (depending on the tokenizer)
  • Examples (using OpenAI’s tokenizer):
    • "Hello, world!" โ†’ ["Hello", ",", " world", "!"] (4 tokens)
    • "I love you" โ†’ ["I", " love", " you"] (3 tokens)
    • A long Chinese sentence โ†’ often 1 Chinese character = 1โ€“2 tokens (less efficient than English)

Why Tokens Matter

  • Context length is measured in tokens (e.g., “this model has an 8K token context”).
  • Cost is usually based on tokens (input tokens + output tokens).
  • Speed depends on how many tokens the model processes.

A Tool is a function you give to an LLM so it can take actions beyond just generating text โ€” it lets the model interact with the real world.

ๅทฅๅ…ท (Tool) ๆ˜ฏไฝ ็ป™ LLM ็š„ไธ€ไธชๅ‡ฝๆ•ฐ๏ผŒ่ฎฉๅฎƒไธๅชๆ˜ฏ็”Ÿๆˆๆ–‡ๅญ—๏ผŒ่€Œๆ˜ฏ่ƒฝ็œŸๆญฃไธŽๅค–็•Œไบคไบ’ใ€ๆ‰ง่กŒๆ“ไฝœใ€‚

Examples:

  • ๐Ÿ” Search (Bing / web) / ๆœ็ดขๅผ•ๆ“Ž
  • ๐Ÿ—„๏ธ Database query (SQL)
  • ๐Ÿ“Š Data processing (Python)
  • ๐Ÿ”— APIs (CRM, ERP)
  • ๐Ÿ“ File reading

Why tools matter?

Because LLM alone:

  • cannot access real-time data
  • cannot query enterprise systems
  • cannot execute actions

๐Ÿ‘‰ Tools = โ€œhands of the modelโ€ / model ็š„โ€œๆ‰‹โ€

Tool Schema

Tool Calling โ€” LLM Decision & Tool Selection

Tool Execution And Result Return

Multi-turn Tool Loop

ReAct Mode Implementation


Workflow Rules = logic that controls how an agent behaves
ๆŽงๅˆถ Agent ่กŒไธบ็š„โ€œๆต็จ‹่ง„ๅˆ™โ€

Examples:

  • Step ordering
  • Tool selection rules
  • Approval conditions
  • Safety constraints

Example

1. Understand intent / ็†่งฃ้—ฎ้ข˜
2. Check memory / ๆŸฅ็œ‹่ฎฐๅฟ†
3. Decide if tool is needed / ๅˆคๆ–ญๆ˜ฏๅฆ่ฆๅทฅๅ…ท
4. Call tool (if needed) / ่ฐƒ็”จๅทฅๅ…ท
5. Combine results / ๆฑ‡ๆ€ป็ป“ๆžœ
6. Generate final answer / ่พ“ๅ‡บ็ญ”ๆกˆ


Appendix

OpenAI Platform Doc – OpenAI Developers

Azure OpenAI Documentations

Summary of AI, ML, LLM

Summary of AI, ML, LLM

AI (Artificial Intelligence) contains ML (Machine Learning), which contains LLM (Large Language Models) focused on languageโ€” like nested Russian dolls.

AI๏ผˆไบบๅทฅๆ™บ่ƒฝ๏ผ‰ๅŒ…ๅซ ML๏ผˆๆœบๅ™จๅญฆไน ๏ผ‰๏ผŒML ๅ†ๅŒ…ๅซ LLM๏ผˆๅคง่ฏญ่จ€ๆจกๅž‹๏ผ‰ไธ“ๅš่ฏญ่จ€็š„ๆจกๅž‹โ€”โ€”ๅฐฑๅƒไฟ„็ฝ—ๆ–ฏๅฅ—ๅจƒไธ€ๆ ทๅฑ‚ๅฑ‚ๅตŒๅฅ—ใ€‚

TermEnglish (One-line Human Explanation)ไธญๆ–‡๏ผˆไธ€ๅฅ่ฏไบบ่ฏ่งฃ้‡Š๏ผ‰
AI (Artificial Intelligence)The broad field of making computers behave intelligently like humans.AI ๆ˜ฏโ€œ่ฎฉ็”ต่„‘ๅƒไบบไธ€ๆ ทไผšๆ€่€ƒใ€ไผšๅšไบ‹โ€็š„ๆ€ป้ข†ๅŸŸใ€‚
ML (Machine Learning)A subset of AI where computers learn patterns from data instead of hard-coded rules.ML ๆ˜ฏ AI ็š„ไธ€็งๆ–นๅผ๏ผš่ฎฉ็”ต่„‘ไปŽๆ•ฐๆฎ้‡Œโ€œ่‡ชๅทฑๅญฆ่ง„ๅพ‹โ€๏ผŒ่€Œไธๆ˜ฏไบบๆ‰‹ๅ†™่ง„ๅˆ™ใ€‚
LLM (Large Language Model)A type of ML model trained on huge amounts of text to understand and generate language.LLM ๆ˜ฏ ML ็š„ไธ€็งๅคงๅž‹่ฏญ่จ€ๆจกๅž‹๏ผŒไธ“้—จๅญฆไน ๆตท้‡ๆ–‡ๅญ—๏ผŒไปŽ่€Œไผšโ€œ่Šๅคฉใ€ๅ†™ไฝœใ€ๅ›ž็ญ”้—ฎ้ข˜โ€ใ€‚

Frequently Used GenAI & LLM Concepts (Simple Explanations)

Agent

An AI Agent is a system that can autonomously break down tasks, make decisions, and execute actions using tools and reasoning.

Agentic Workflow

An agentic workflow is a multi-step autonomous process where an AI system completes tasks without continuous human intervention.
or says: AI auto-complete entire process without human intervention.
e.g.
apply for –> validation –> calculation –> output results
human intervention.

Chunking

Splitting big documents into small pieces so AI can handle them better.

e.g. There are 200 pages in a PDF file, AI cannot read all at once, so splitting file into many small pieces/chunks.
1st piece: 1- 500 words;
2nd piece/chunk: 501 – 1000 words;
3rd piece/chunk: 1001 – 1500 words;
…..
each chunk will become embedding.
It is commonly used in RAG systems to prepare documents for embedding and retrieval.

Cosine Similarity

Cosine similarity measures how similar two vectors are in meaning by comparing their direction in vector space.
or says: A way to measure how similar two pieces of meaning are.
e.g.
apple vs banana : yes, they are very similar.
apple vs car: no, they are not similar at all.

Context Window

Context window is the maximum amount of text an LLM can process at once.

Embedding

Embeddings convert text into numerical vectors that represent meaning.

โ€œappleโ€ become [0.12, -0.98, 0.33, ……]
โ€œorangeโ€ become [0.12, -0.98, 0.456, ,,,,,,,] too,
so AI will find
apple = Fruit,
apple != car
or says “simile to a fruit”, and it is not a car. Similar meanings result in closer vector distances, allowing machines to compare semantic similarity instead of exact words.

Fine-tuning

Fine-tuning is the process of further training a pre-trained model on domain-specific data to improve performance in a specialized area.

Hallucination

Hallucination occurs when an LLM generates incorrect or fabricated information while sounding confident.

LangChain

LangChain is a framework for building applications powered by LLMs by connecting models with tools, APIs, and data sources.
in short, chaining interlinkage/link AI , Data, Tools …….
or says “A tool to connect LLMs, data, and tools into applications.”

LangGraph

LangGraph is a framework for building stateful, graph-based AI workflows where agents can loop, branch, and maintain memory across steps.
or says: A workflow system that lets AI follow multi-step flows with loops and decisions.
e.g. SQL Agent,
Write SQL script –> Execute –> Error Alert –> Fix –> Re-try

LLM

Large Language Model. The AI brain that can understand and generate language.

e.g. user asks AI “please write an email”, then output a completed email.
Company uses it to generate Report, analyst Data, auto reply client, …..

MCP

MCP (Model Context Protocol) defines a standardized way for LLMs to interact with external tools, APIs, and data systems.
or says: A standard way for AI to use tools and data systems.

Model Drift

Model drift occurs when a deployed modelโ€™s performance degrades due to changes in real-world data over time.
or says: After the AI was put into use, it started to make mistakes.
why/what’s happened?
maybe, training used old data, now data has changed/updated.

Prompt

A prompt is the instruction given to an LLM.
Well-designed prompts significantly improve the quality and accuracy of model outputs.
e.g.
bad prompt: “write a letter”, — not clearly, what letter you need, thank you letter? complaining letter? ,,,,
good prompt: “Please write a thank you letter to Mary since she gave me a gift.”

Prompt Engineering

Prompt engineering is the practice of designing effective prompts to guide LLM behavior and improve output quality.
or say: Designing better instructions to improve AI responses.

RAG

Retrieval-Augmented Generation. Retrieval-Augmented Generation combines retrieval and generation.
The system first retrieves relevant documents, then uses an LLM to generate an answer based on that information.
e.g. look up HR documents –> pass documents to GPT –> GPT summary then answer question.

Retrieval

Retrieval is the process of searching a knowledge base or vector database to find relevant information before generating an answer.
e.g.
“what is the return policy?” , AI system will look up in “vector DB”, find out “policy document”, pass it to GPT, then answer the question – what is the return policy?

Token

Tokens are the smallest units of text that an LLM processes.

e.g. a sentence like โ€œI love Torontoโ€, AI splits โ€œI love Torontoโ€ into smaller pieces before the model can understand it.

  • I
  • love
  • Toronto

these are tokens,
Token count also determines cost and context limits in LLM systems.

Tool Calling

Tool calling allows LLMs to execute external functions such as APIs, databases, or code to perform real-world actions.
e.g. “AI can “take action”.
>search order
> search database

Vector DB

A database that stores meaning-based vectors for similarity search. It allows AI systems to retrieve semantically relevant documents instead of keyword-based search.
e.g.
there are 10000 file,
HR policy,
IT manual,
Finance report,
……
all of those files become embedding saved in Vector DB. When user asks question, AI will not use “Key-words” to seek, it uses “mean” to match.

Vector Search

Vector search retrieves results based on semantic similarity rather than keyword matching.
or says: Searching by meaning instead of exact words.