<oai_dc:dc xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
  <dc:contributor>Mira, Antonietta</dc:contributor>
  <dc:creator>Denti, Francesco</dc:creator>
  <dc:date>2020-02-07</dc:date>
  <dc:description xmlns:ns0="xml" ns0:lang="en">Bayesian mixture models are ubiquitous in statistics due to their simplicity and flexibility and can be easily employed in a wide variety  of contexts. In this dissertation, we aim at providing a few contributions to current Bayesian data analysis methods, often motivated  by research questions from biological applications. In particular, we focus on the development of novel Bayesian mixture models,  typically in a nonparametric setting, to improve and extend active research areas that involve large-scale data: the modeling of  nested data, multiple hypothesis testing, and dimensionality reduction. Therefore, our goal is twofold: to develop robust statistical  methods motivated by a solid theoretical background, and to propose efficient, scalable and tractable algorithms for their  applications. The thesis is organized as follows. In Chapter 1 we shortly review the methodological background and discuss the  necessary concepts that belong to the different areas that we will contribute to with this dissertation. In Chapter 2 we propose a  Common Atoms model (CAM) for nested datasets, which overcomes the limitations of the nested Dirichlet Process, as discussed in  Camerlenghi et al.,2018. We derive its theoretical properties and develop a slice sampler for nested data to obtain an efficient  algorithm for posterior simulation. We then embed the model in a Rounded Mixture of Gaussian kernels framework to apply our  method to an abundance table from a microbiome study. In Chapter 3 we develop a BNP version of the two-group model (Efron,  2004), modeling both the null density f_0 and the alternative density f_1 with Pitman-Yor process mixture models. We propose to fix  the two discount parameters sigma_0 and sigma_1 so that sigma_0&gt;sigma_1, according to the rationale that the null PY should be  closer to its base measure (appropriately chosen to be a standard Gaussian base measure), while the alternative PY should have  fewer constraints. To induce separation, we employ a non-local prior (Johnson and Rossell, 2010) on the location parameter of the  base measure of the PY placed on f_1. We show how the model performs in different scenarios and apply this methodology to a  microbiome dataset. Chapter 4 presents a second proposal for the two-group model. Here, we make use of non-local distributions to  model the alternative density directly in the likelihood formulation. We propose both a parametric and a nonparametric formulation of  the model. We provide a theoretical justification for the adoption of this approach and, after comparing the performance of our model  with several competitors, we present three applications on real, publicly available genomic datasets. In Chapter 5 we focus on  improving the model for intrinsic dimensions (IDs) estimation discussed in Allegra et al.,2019. In particular, the authors estimate the  IDs modeling the ratio of the distances from a point to its first and second nearest neighbors (NNs). First, we propose to include more  suitable priors in their parametric, finite mixture model. Then, we extend the existing theoretical methodology by deriving closed-form  distributions for the ratios of distances from a point to two NNs of generic order. We propose a simple Dirichlet process mixture  model, where we exploit the novel theoretical results to extract more information from the data. The chapter is then concluded with  simulation studies and the application to real data. Finally, Chapter 6 presents the future directions and concludes.</dc:description>
  <dc:description xmlns:ns1="xml" ns1:lang="it">I modelli mistura bayesiani sono onnipresenti in statistica per la loro semplicità e flessibilità e possono essere facilmente  impiegati in un'ampia varietà di contesti. In questa tesi, miriamo a fornire alcuni contributi agli attuali metodi bayesiani di  analisi dei dati, spesso motivati ​​da domande di ricerca provenienti da applicazioni biologiche. In particolare, ci concentriamo  sullo sviluppo di nuovi modelli mistura bayesiani, tipicamente in un ambiente non parametrico, per migliorare ed estendere  aree di ricerca che coinvolgono dati caratterizzati da grande dimensioni: la modellazione di dati nested, test di ipotesi  simultaneo e la riduzione della dimensionalità. Pertanto, il nostro obiettivo è duplice: sviluppare metodi statistici robusti  motivati da un solido background teorico e proporre algoritmi efficienti, scalabili e trattabili per le loro applicazioni. La tesi è  organizzata come segue. Nel capitolo 1 esamineremo brevemente il background metodologico e discuteremo i concetti  necessari che appartengono alle diverse aree a cui contribuiremo con questa tesi. Nel capitolo 2 proponiamo un modello di  atomi comuni (CAM) per nested data, che supera le limitazioni del processo del nested Dirichlet Process, come discusso in  Camerlenghi et al.,2018. Deriviamo le sue proprietà teoriche e sviluppiamo uno slice sampler per dati nested al fine di  ottenere un algoritmo efficiente per la simulazione della posterior. Abbiamo poi incorporato il modello in un framework di  Rounded mixture of Gaussian Kernels, così da applicare il nostro metodo a una abundance table derivante da uno studio di  microbioma. Nel capitolo 3 sviluppiamo una versione BNP del two-group model, modellando sia f_0 che f_1 con Pitman-Yor  mixtures models. Proponiamo di fissare i due parametri sigma_0 e sigma_1 in modo che sigma_0&gt;sigma_1, in base alla  logica secondo cui il PY che modella la distribuzione nulla dovrebbe essere più vicino alla sua misura di base  (opportunamente scelta Gaussiana standard), mentre il PY alternativo dovrebbe avere meno vincoli. Per indurre la  separazione, impieghiamo una non-local prior (Johnson and Rossell, 2010) sul parametro location della misura base del PY  collocato su f_1. Mostriamo come il modello si comporta in diversi scenari e applichiamo questa metodologia a un set di dati  del microbioma. Il capitolo 4 presenta una seconda proposta per il two-group model. Qui, utilizziamo non-local distributions  per modellare la densità alternativa direttamente nella formulazione della Likelihood. Abbiamo proposto una formulazione  sia parametrica che non parametrica del modello. Forniamo poi una giustificazione teorica per l'adozione di questo  approccio e, dopo aver confrontato le prestazioni del nostro modello con diversi concorrenti, presentiamo tre applicazioni su  set di dati genomici reali pubblicamente disponibili. Nel capitolo 5 ci concentriamo sul miglioramento del modello per la stima  delle dimensioni intrinseche (ID) discusso in Allegra et al.,2019, dove gli autori stimano gli IDs modellando il rapporto delle  distanze da un punto dal suo primo e secondo vicino più vicino (NN). Innanzitutto, proponiamo di includere distribuzioni a  priori più adatte nel loro modello mistura finita. Quindi, estendiamo la metodologia teorica esistente derivando distribuzioni in  forma chiusa per i rapporti di distanze da un punto a due NNs di ordine generico. Proponiamo poi un semplice modello di  mistura nonparametrica usando il processo di Dirichlet, in cui sfruttiamo le distribuzioni derivate per estrarre più informazioni  dai dati. Il capitolo si conclude quindi con studi di simulazione e l'applicazione a dati reali. Infine, il capitolo 6 presenta le  direzioni future e le conclusioni.</dc:description>
  <dc:format>application/pdf</dc:format>
  <dc:identifier>https://susi.usi.ch/global/documents/319157</dc:identifier>
  <dc:identifier>https://n2t.net/ark:/12658/srd1319157</dc:identifier>
  <dc:identifier>https://susi.usi.ch/documents/319157/files/2020ECO001.pdf</dc:identifier>
  <dc:language>eng</dc:language>
  <dc:relation>info:eu-repo/semantics/altIdentifier/urn/urn:nbn:ch:rero-006-120584</dc:relation>
  <dc:relation>info:eu-repo/semantics/altIdentifier/ark/12658/srd1319157</dc:relation>
  <dc:rights>info:eu-repo/semantics/openAccess</dc:rights>
  <dc:rights>License undefined</dc:rights>
  <dc:subject xmlns:ns2="xml" ns2:lang="en">Bayesian data analysis</dc:subject>
  <dc:subject xmlns:ns3="xml" ns3:lang="en">Bayesian mixture model</dc:subject>
  <dc:subject xmlns:ns4="xml" ns4:lang="en">Bayesian nonparametrics</dc:subject>
  <dc:subject xmlns:ns5="xml" ns5:lang="en">Dirichlet process</dc:subject>
  <dc:subject xmlns:ns6="xml" ns6:lang="en">Multiple hypothesis testing</dc:subject>
  <dc:subject xmlns:ns7="xml" ns7:lang="en">Intrinsic dimension</dc:subject>
  <dc:subject xmlns:ns8="xml" ns8:lang="en">Poisson process</dc:subject>
  <dc:subject xmlns:ns9="xml" ns9:lang="en">Nested data</dc:subject>
  <dc:subject xmlns:ns10="xml" ns10:lang="en">Partial exchangeability</dc:subject>
  <dc:subject xmlns:ns11="xml" ns11:lang="en">Common atom model</dc:subject>
  <dc:subject xmlns:ns12="xml" ns12:lang="en">Non-local distributions</dc:subject>
  <dc:subject xmlns:ns13="xml" ns13:lang="it">Analisi dei dati Bayesian</dc:subject>
  <dc:subject xmlns:ns14="xml" ns14:lang="it">Modelli mistura bayesiani</dc:subject>
  <dc:subject xmlns:ns15="xml" ns15:lang="it">Bayesiana nonparametrica</dc:subject>
  <dc:subject xmlns:ns16="xml" ns16:lang="it">Processo di Dirichlet</dc:subject>
  <dc:subject xmlns:ns17="xml" ns17:lang="it">Test di ipotesi multiple</dc:subject>
  <dc:subject xmlns:ns18="xml" ns18:lang="it">Dimensione intrinseca</dc:subject>
  <dc:subject xmlns:ns19="xml" ns19:lang="it">Processo di Poisson</dc:subject>
  <dc:subject xmlns:ns20="xml" ns20:lang="it">Dati annidati</dc:subject>
  <dc:subject xmlns:ns21="xml" ns21:lang="it">Scambiabilità parziale</dc:subject>
  <dc:subject xmlns:ns22="xml" ns22:lang="it">Modello agli atomi comuni</dc:subject>
  <dc:subject xmlns:ns23="xml" ns23:lang="it">Distribuzioni non-locali</dc:subject>
  <dc:subject>info:eu-repo/classification/udc/33</dc:subject>
  <dc:title xmlns:ns24="xml" ns24:lang="en">Bayesian mixtures for large scale inference</dc:title>
  <dc:type>http://purl.org/coar/resource_type/c_db06</dc:type>
</oai_dc:dc>
