<oai_dc:dc xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
  <dc:contributor>Crestani, Fabio</dc:contributor>
  <dc:creator>Inches, Giacomo</dc:creator>
  <dc:date>2014-01-29</dc:date>
  <dc:description xmlns:ns0="xml" ns0:lang="en">In recent years short user-generated documents have been gaining popularity on the  Internet and attention in the research communities. This kind of documents are generated  by users of the various online services: platforms for instant messaging communication,  for real-time status posting, for discussing and for writing reviews. Each of these services  allows users to generate written texts with particular properties and which might require  specific algorithms for being analysed. In this dissertation we are presenting our work  which aims at analysing this kind of documents. We conducted qualitative and quantitative  studies to identify the properties that might allow for characterising them. We compared  the properties of these documents with the properties of standard documents employed in  the literature, such as newspaper articles, and defined a set of characteristics that are  distinctive of the documents generated online. We also observed two classes within the  online user-generated documents: the conversational documents and those involving  group discussions. We later focused on the class of conversational documents, that are  short and spontaneous. We created a novel collection of real conversational documents  retrieved online (e.g. Internet Relay Chat) and distributed it as part of an international  competition (PAN @ CLEF'12). The competition was about author characterisation, which  is one of the possible studies of authorship attribution documented in the literature. Another  field of study is authorship identification, that became our main topic of research. We  approached the authorship identification problem in its closed-class variant. For each  problem we employed documents from the collection we released and from a collection of  Twitter messages, as representative of conversational or short user-generated  documents. We proved the unsuitability of standard authorship identification techniques for  conversational documents and proposed novel methods capable of reaching better  accuracy rates. As opposed to standard methods that worked well only for few authors,  the proposed technique allowed for reaching significant results even for hundreds of  users.</dc:description>
  <dc:format>application/pdf</dc:format>
  <dc:identifier>https://susi.usi.ch/global/documents/318463</dc:identifier>
  <dc:identifier>https://n2t.net/ark:/12658/srd1318463</dc:identifier>
  <dc:identifier>https://susi.usi.ch/documents/318463/files/2014INFO019.pdf</dc:identifier>
  <dc:language>eng</dc:language>
  <dc:relation>info:eu-repo/semantics/altIdentifier/urn/urn:nbn:ch:rero-006-114051</dc:relation>
  <dc:relation>info:eu-repo/semantics/altIdentifier/ark/12658/srd1318463</dc:relation>
  <dc:rights>info:eu-repo/semantics/openAccess</dc:rights>
  <dc:rights>License undefined</dc:rights>
  <dc:subject xmlns:ns1="xml" ns1:lang="en">Short user-generated documents</dc:subject>
  <dc:subject xmlns:ns2="xml" ns2:lang="en">Instant messaging</dc:subject>
  <dc:subject xmlns:ns3="xml" ns3:lang="en">Chat</dc:subject>
  <dc:subject xmlns:ns4="xml" ns4:lang="en">Forum</dc:subject>
  <dc:subject xmlns:ns5="xml" ns5:lang="en">Blog</dc:subject>
  <dc:subject xmlns:ns6="xml" ns6:lang="en">Online user-generated documents</dc:subject>
  <dc:subject xmlns:ns7="xml" ns7:lang="en">Conversational documents</dc:subject>
  <dc:subject xmlns:ns8="xml" ns8:lang="en">Authorship attribution</dc:subject>
  <dc:subject xmlns:ns9="xml" ns9:lang="en">Authorship identification</dc:subject>
  <dc:subject xmlns:ns10="xml" ns10:lang="en">Twitter</dc:subject>
  <dc:subject xmlns:ns11="xml" ns11:lang="en">Mrr</dc:subject>
  <dc:subject xmlns:ns12="xml" ns12:lang="en">Information retrieval</dc:subject>
  <dc:subject xmlns:ns13="xml" ns13:lang="en">Text mining</dc:subject>
  <dc:subject xmlns:ns14="xml" ns14:lang="en">Burstiness</dc:subject>
  <dc:subject xmlns:ns15="xml" ns15:lang="en">Interlocutors</dc:subject>
  <dc:subject>info:eu-repo/classification/udc/004</dc:subject>
  <dc:title xmlns:ns16="xml" ns16:lang="en">Statistical models for the analysis of short user-generated documents : author identification for conversational documents</dc:title>
  <dc:type>http://purl.org/coar/resource_type/c_db06</dc:type>
</oai_dc:dc>
