Top X 2 0 A Very Fast ObjectStore
Top. X 2. 0 — A (Very) Fast Object-Store for Top-k XPath Query Processing Martin Theobald Stanford University Mohammed Abu. Jarour Hasso-Plattner Institute Ralf Schenkel Max-Planck Institute
//article[. //bib[about(. //item, “W 3 C”)] ]//sec[about(. //, “XML retrieval”)] //par[about(. //, “native XML databases”)] RANKING article title “Current Approaches to XML Data Manage- sec ment” XML Files” sec VAGUENESS bib par “XML queries with an expressive power similar to that of Datalog …” title “Native XML Data Bases. ” “Native XML data base systems can store schemaless data. . . ” par “Data management systems control data acquisition, storage, and retrieval. Systems evolved from flat files … ” title sec title par title “The item “The Ontology Game” bib title “The Dirty Little Secret” item title EARLY PRUNING par inproc “XML-QL: “Proc. Query A Query Languages Language Workshop, for XML. ” W 3 C, 1998. ” par “Sophisticated technologies “There, I've said developed by it - the "O" word. If smart people. ” anyone is thinking along ontology lines, I would like to break some old news …” “XML” par url “w 3 c. org/xml” “What does XML add for retrieval? It adds formal ways …” From the INEX ’ 03 -’ 05 IEEE Collection
Frontends • Web Interface • Web Service • API Non-conjunctive Top-k XPath Query Processing Probabilistic Candidate Pruning SA SA SA Scan Threads Dynamic Query Expansion Ontology/ Large Thesaurus Word. Net, Open. Cyc, etc. Random Access Candidate Queue Sequential Access Probabilistic Index Access Scheduling Top. X 1. 0 2. 0 Query Processor Top-k Queue Expensive Predicates • Path Conditions • Phrases & Proximity • Other Full-Text Op’s Index Metadata • Selectivities • Histograms • Correlations RA JDBC Relational DBMS Backend Unified Text & XML Schema Indexer/Crawler RA
Data Model ftf (“xml”, article 1 ) = 4 “xml data manage xml manage system vary “xml data manage xml manage system wide expressive power native vary wide expressive powerxml native xml data base system data base native xmlstore data base system storeschemalessdata“ <article> <title> XML Data Management </title> <abs>XML management systems vary widely in their expressive power. </abs> <sec> <title>Native XML Data Bases. </title> <par>Native XML data base systems can store schemaless data. </par> </sec> </article> article 1 6 title 2 1 abs 2 3 “xml data “xml manage” system vary wide expressive power“ “native xml xmldatabase system store schemaless data“ sec 4 5 title 5 3 “native xml data base” ftf (“xml”, sec 4 ) = 2 Ø XML trees (no XLink/ID/IDRef) Ø Pre-/postorder ranges for the structural index Ø Redundant full-content text nodes 6 par 4 “native xml data base system store schemaless data“
Scoring Model [INEX ‘ 05/’ 06/’ 07] Content Index (Tag-Term Pairs) Element Freq. Element Statistics Doc. ID Tag Term Pre Post FTF Tag Term EF Tag N Av. Len 1 article xml 1 6 4 article xml 863 article 659 K 269. 2 1 sec xml 4 5 2 sec xml 947 sec 1. 6 M 89. 1 1 title xml 5 3 1 title xml 62 title 2. 2 M 2. 8 1 par xml 6 4 1 par xml 674 par 2. 8 M 34. 1 … … … Ø XML-specific variant of Okapi BM 25 bib[“transactions”] vs. par[“transactions”] (originating from probabilistic IR on unstructured text)
Top. X 1. 0: Relational Schema Ø Precompute & materialize scoring model into combined inverted index over tag-term pairs Ø Supports sorted access (by Max. Score) and random access (by Doc. ID) sec[“xml”] Doc. ID RA SA Pre Post Score Max Score 2 2 15 0. 9 2 10 8 0. 5 0. 9 1 23 48 0. 8 1 45 87 0. 2 0. 8 3 4 24 0. 7 3 12 18 0. 6 0. 7 3 17 25 0. 3 0. 7 … … … Two B+trees Select Doc. ID, Pre, Post, Score From Tag. Term. Index Where tag=‘sec’ and term=‘xml’ Order by Max. Score desc, Doc. ID desc Pre asc, Post Desc Select Pre, Post, Score From Tag. Term. Index Where Doc. ID=3 and tag=‘sec’ and term=‘xml’ Order by Pre Asc, Post Desc
Top-k XPath on a Relational Schema [VLDB ’ 05] • Content-only (CO) & “structure enriched” queries: //sec[about(. //, “XML”) and about(. //title, “native”]//par[about(. //, “retrieval”)] sec[“xml”] title[“native”] par[“retrieval”] Doc. ID Pre Post Score Max Score 2 2 15 0. 9 17 2 15 0. 9 1 12 21 1. 0 2 10 8 0. 9 3 14 10 0. 8 0. 9 2 8 14 0. 8 1 23 48 0. 8 2 4 12 0. 5 5 3 7 1 0. 7 1 45 87 0. 2 0. 8 31 12 23 0. 4 4 6 4 1 0. 7 … … … … Ø Ø Ø Sequentially (mostly) scan each index list in desc. order of Max. Score Hash-join element blocks by Doc. ID in-memory Do “some” incremental XPath evaluation using Pre/Post indices Aggregate Score along connected path fragments Use variant of Fagin’s threshold algorithm for top-k-style early termination
Top-k XPath on a Relational Schema [VLDB ’ 05] • Content-and-structure (CAS) queries: sec[“xml”] SA //article//sec[about(. //, “XML”)] 1. 0 article Doc ID Pre Post 0. 9 1 1 245 0. 8 0. 9 2 1 123 48 0. 8 3 1 176 45 87 0. 2 0. 8 4 1 89 … … … Doc. ID Pre Post Score Max Score 2 2 15 0. 9 2 10 8 1 23 1 … RA Ø Expensive predicate probes (RA) to the structure index (3 rd B+tree) Select Pre, Post From Tag. Index Where Doc. ID=2123 and Tag=‘article’ Order by Pre asc, Post desc Ø Non-conjunctive XPath evaluations Ø Dynamically relax content- & structure-related query conditions (top-k results entirely driven by score aggregations for content & structure cond. ’s)
Relational Schema (cont’d) Ø No shredding into DTD-specific relational schema! Ø No DTD at all for INEX Wikipedia! sec[“xml”] article Doc ID Pre Post Score Max Score Doc ID Pre Post 2 2 15 0. 9 1 1 245 2 10 8 0. 9 2 1 123 1 23 48 0. 8 3 1 176 1 45 87 0. 2 0. 8 4 1 89 … … … … 20, 810, 942 distinct tag-term pairs for 4. 38 GB Wikipedia collection 1, 107 distinct tags
Relational Schema (cont’d) Content Index Doc. I D Tag Term 2 sec 2 Structure Index Pre Post Score Max Score Doc ID Tag Pre Post xml 2 15 0. 9 1 article 1 245 sec xml 10 8 0. 5 0. 9 2 article 1 123 17 title xml 5 3 0. 5 3 sec 2 15 1 par xml 6 4 0. 7 3 sec 10 8 … … … … … (4+4+4+4) bytes X 567, 262, 445 tag-term pairs … (4+4+4+4) bytes X 52, 561, 559 tags 16 GB 0. 85 GB Ø 2 -dimensional source of redundancy Ø Full-content scoring model (#terms times avg. depth of a text node 6. 7 for INEX Wiki) Ø De-normalized relational schema Ø High overhead in the architecture (Java->JDBC->DBMS & back) Ø Element-block sizes are data-driven, not easy to control layout on disk Ø Hashing too slow compared to very efficient in-memory merge-joins
Top. X 2. 0: Object-Oriented Storage Binary file 2 Doc. ID 2 15 0. 9 10 8 0. 5 B L 1 0 Max. Sore 23 48 0. 8 Max. Sore 45 87 0. 2 title[“xml”] 122, 564 17 2 15 0. 9 14 5 0. 2 27 32 0. 4 Tag Term Pre Post Score Max Score 2 sec xml 2 15 0. 9 2 sec xml 10 8 0. 5 0. 9 … … … … 17 title xml 2 15 0. 9 … … … … 11 par xml 6 4 0. 7 … … … … (4+4+4+4) X 567, 262, 445 Relational: 16 GB 3 B L sec[“xml”] Doc. I D 1 6 0. 9 par[“xml”] … B – Element block separator L – Index list separator 432, 534 + 4 X 456, 466, 649 (4+4+4) X 567, 262, 445 Object-oriented: 8. 6 GB (+ (4+4) X 20, 810, 942 = 166 MB for the offset index)
Object-Oriented Storage w/Block-Merging Document Block sec[“xml”] 0 1 B B 23 48 Ø Group element blocks with similar Max. Score into document blocks of fixed length (e. g. 256 KB) 0. 8 2 2 15 0. 9 10 8 0. 5 Max. Sore 5 2 24 0. 7 3 11 0. 3 Ø Sort element blocks within each document block by Doc. ID B B…B 3 5 23 0. 5 7 21 0. 3 24 15 0. 1 B 6 6 15 0. 6 13 17 0. 5 14 32 0. 3 Ø Supports Ø Sorted access by Max. Score Ø Merge-joins by Doc. ID Max. Sore Ø Raw disk access B B…B L … … title[“xml”] 122, 564
Merging Document Blocks sec[“xml”] 0. 8 SA par[“retrieval”] 2 1 B B 0. 7 //sec[about(. //, “XML”)] //par[about(. //, “retrieval”)] 23 48 0. 8 2 B 12 48 5 15 0. 9 3 17 0. 9 10 8 0. 5 13 9 0. 2 B 7 2 24 0. 7 65 21 1. 0 3 11 0. 3 72 43 0. 5 B B…B 3 6 5 23 0. 6 18 29 0. 8 7 21 0. 3 23 24 0. 8 24 15 0. 1 24 15 0. 7 B 6 B 0. 8 9 6 15 0. 5 32 45 0. 8 13 17 0. 5 33 27 0. 7 14 32 0. 3 37 39 0. 5 B B…B … 0. 9 2 5 1. 0 B B…B … Sequential access and efficient merge-joins on top of large document blocks
Compressed Number Encoding Ø Multi-attribute (4), double-nested block-index structure Ø Delta encoding only works for Doc. ID (and to some extent for Pre) Ø No specific assumptions on distributions of Pre/Post or Score Ø No Unary or Huffman coding (prefix-free but additional coding table) Ø Sophisticated compression schemes may be expensive to decode Ø No Zip, etc. Ø But known number ranges Ø Doc. ID [1, 659, 388] -> 3 bytes (2543 = 16, 387, 064, lossless) Ø Pre/Post [1, 43, 114] -> 2 bytes (2542 = 64, 516, lossless) Ø Score [0, 1] -> rounded to 1 byte (254 buckets, lossy) Variable-length byte encoding w/leading length-indicator byte Len Pre Post Score 4 3 26 7 9 5 bytes 225 332 192 10 bytes
Some more tricks… 0 score EB k … DB 2 (256 KB) DBl (256 KB) … … 1 EB 2 freq DB 1 (256 KB) EB 1 sec[“xml”] Histogram Block 36 bytes Ø Dump leading histogram blocks into index list headers Ø Histograms only for index lists that exceed one document block (<5% of all lists) Ø Own native compare methods for Doc. ID, Pre/Post Ø Decode only Score for arithmetic op’s ( Mostly perform pointer operations at qp time) Incrementally read & process precomputed memory image for fast top-k queries on top of large disk blocks
Block Access Scheduling [VLDB ’ 06] Ø SA Scheduling Inverted Block-Index (256 KB doc-blocks) 0. 9 1. 0 0. 9 Δ 3, 3 = 0. 2 0. 8 Ø RA Scheduling Δ 1, 3 = 0. 8 0. 7 Ø 2 -phase probing: 0. 2 … … 0. 6 … 0. 9 1. 0 SA SA SA 1. 0 Ø Look-ahead Δi through precomputed score histograms Ø Knapsack-based optimization of Score Reduction 0. 8 Schedule RAs “late & last” RA i. e. , cleanup the queue if Ø Extended probabilistic cost model for integrated SA & RA scheduling
Object Storage Summary (incl. histograms) Relational 12 Object-Oriented 8 4 0 Structure Index Content Index 16 • • 567, 262, 445 tag-term pairs 20, 810, 942 distinct tag-term pairs 20, 815, 884 document blocks (256 KB) 456, 466, 649 element blocks • • 52, 561, 559 tags (elements) 1, 107 distinct tags 2, 323 document blocks (256 KB) 8, 999, 193 element blocks • 4, 703, 385, 686 total bytes • 246, 601, 752 total bytes (8. 3 bytes/tag-term pair) (4. 7 bytes/tag) 4. 38 GB Wikipedia XML sources
Preliminary Runtime Experiments CO (top-10, non-conjunctive)
Preliminary Runtime Experiments CAS (top-10, non-conjunctive)
Some INEX Results CAS (top-1, 500, non-conjunctive)
Some INEX Results CAS (top-1, 500, non-conjunctive)
Conclusions & Outlook Ø Scalable and efficient XML-IR with vague search Ø Mature system, reference engine for INEX topic development & interactive tracks [VLDB Special Issue on DB&IR Integration ‘ 08] Ø Brand-new Top. X 2. 0 prototype Ø Very efficient reimplementation in C++ Ø Object-oriented XML storage, moderate compression rates Ø 10— 20 times better sequential throughput than relational Ø The Future: More features! Ø Ø Generalized proximity search, graph top-k Updates (gaps within document blocks) XQuery Full-Text (top-k-style bounds over IF, For-Let) …
http: //www. inex. otago. ac. nz/
- Slides: 23