<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://choishio.github.io//feed.xml" rel="self" type="application/atom+xml" /><link href="https://choishio.github.io//" rel="alternate" type="text/html" /><updated>2026-06-30T17:57:57+09:00</updated><id>https://choishio.github.io//feed.xml</id><title type="html">Welcome to SHIO-LAB :)</title><subtitle>This page will be filled with my study &amp; research</subtitle><author><name>Dayun Choi</name></author><entry><title type="html">My research topic has just been decided!</title><link href="https://choishio.github.io//research/research-topic/" rel="alternate" type="text/html" title="My research topic has just been decided!" /><published>2022-06-28T14:00:00+09:00</published><updated>2022-06-28T21:30:00+09:00</updated><id>https://choishio.github.io//research/research-topic</id><content type="html" xml:base="https://choishio.github.io//research/research-topic/"><![CDATA[<p>I have just decided my research topic!</p>

<p>It is “Sound Source Separation based on Audio-Visual Multimodal Learning”.</p>

<blockquote>
  <p>Let’s assume that we are the audiences of an orchestra and we can hear the sounds of violin and flute but cannot hear the sounds of piano well. At this time, we can see the performers playing three instruments respectively. Then we can understand that there are three instruments and the magnitude of piano sounds is little. So the visual information can be used to help to know what sounds there are and separate them.</p>
</blockquote>

<p>There are the following things to implement for this task:</p>
<ul>
  <li>Detection of the objects in video for instance level</li>
  <li>Classification of the sound of them in audio</li>
  <li>Learning the audio-visual model jointly for separation of the sound sources</li>
</ul>

<p>You can refer to the following figures to understand this task.</p>

<p align="center">
    <img src="https://user-images.githubusercontent.com/74304696/176095718-f05e079c-a2b6-45f9-b080-0048c6d922c9.png" />
    Michelsanti, Daniel, et al. "An overview of deep-learning-based audio-visual speech enhancement and separation." IEEE/ACM Transactions on Audio, Speech, and Language Processing 29 (2021): 1368-1396.
</p>

<p align="center">
    <img src="https://user-images.githubusercontent.com/74304696/176095723-69c8ef14-a156-41ae-93d5-a0b3f3196693.png" />
    <img src="https://user-images.githubusercontent.com/74304696/176095726-a38700d7-e957-4dd8-8458-5ba4dc264018.png" />
    Zhao, Hang, et al. "The sound of pixels." Proceedings of the European conference on computer vision (ECCV). 2018.
</p>

<p>I hope that my reserch will be well~ :)</p>]]></content><author><name>DaYun Choi</name></author><category term="Research" /><category term="audio" /><category term="visual" /><category term="separation" /><category term="detection" /><category term="classification" /><category term="extraction" /><summary type="html"><![CDATA[I have just decided my research topic!]]></summary></entry><entry><title type="html">First post</title><link href="https://choishio.github.io//first-post/" rel="alternate" type="text/html" title="First post" /><published>2022-06-08T20:40:00+09:00</published><updated>2022-06-21T14:30:00+09:00</updated><id>https://choishio.github.io//first-post</id><content type="html" xml:base="https://choishio.github.io//first-post/"><![CDATA[<p>I will sometimes upload my study &amp; research (or diary).</p>

<p>Now, there is no content since I have just made this page.</p>

<p>So please wait until the frame of this website is completed!</p>]]></content><author><name>DaYun Choi</name></author><summary type="html"><![CDATA[I will sometimes upload my study &amp; research (or diary).]]></summary></entry></feed>