<oai_dc:dc xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
  <dc:creator>Schiavio, Filippo</dc:creator>
  <dc:creator>Bonetta, Daniele</dc:creator>
  <dc:creator>Binder, Walter</dc:creator>
  <dc:date>2020</dc:date>
  <dc:description xmlns:ns0="xml" ns0:lang="en">Big-data systems have gained significant momentum, and Apache Spark is becoming a de-facto standard for modern data analytics. Spark relies on SQL query compilation to optimize the execution performance of analytical workloads on a variety of data sources. Despite its scalable architecture, Spark's SQL code generation suffers from significant runtime overheads related to data access and de-serialization. Such performance penalty can be significant, especially when applications operate on human-readable data formats such as CSV or JSON. In this paper we present a new approach to query compilation that overcomes these limitations by relying on run-time profiling and dynamic code generation. Our new SQL compiler for Spark produces highly-efficient machine code, leading to speedups of up to 4.4x on the TPC-H benchmark with textual-form data formats such as CSV or JSON.</dc:description>
  <dc:format>application/pdf</dc:format>
  <dc:identifier>https://susi.usi.ch/global/documents/322408</dc:identifier>
  <dc:identifier>https://n2t.net/ark:/12658/srd1322408</dc:identifier>
  <dc:identifier>https://susi.usi.ch/documents/322408/files/Schiavio_2020_vldb.pdf</dc:identifier>
  <dc:language>eng</dc:language>
  <dc:relation>https://doi.org/10.14778/3377369.3377382</dc:relation>
  <dc:relation>info:eu-repo/semantics/altIdentifier/ark/12658/srd1322408</dc:relation>
  <dc:rights>info:eu-repo/semantics/openAccess</dc:rights>
  <dc:rights>CC BY-NC-ND</dc:rights>
  <dc:source>Proceedings of the VLDB Endowment. - 2020, vol. 13, no. 5, p. 754-767</dc:source>
  <dc:subject>info:eu-repo/classification/udc/004</dc:subject>
  <dc:title xmlns:ns1="xml" ns1:lang="en">Dynamic speculative optimizations for SQL compilation in Apache Spark</dc:title>
  <dc:type>http://purl.org/coar/resource_type/c_6501</dc:type>
</oai_dc:dc>
