Preskoči na sadržaj
Matematički fakultet

Akademska 2022/23. godina

8 sastanaka, od najnovijeg

Ponedeljak, 10. jul 2023. u 19 časova, na platformi webex computing.math.rs/meet

dr Slavko Gajin
Detekcija anomalija u strukturi mrežnog saobraćaja

Apstrakt

U izlaganju se predstavljaju rezultati istraživanja, razvoja i implementacije sistema za detekciju anomalija u strukturi mrežnog saobraćaja. Imajući u vidu nove vrste napada (eng. zero day attacks) i sve veće korišćenje šifrovanog mrežnog saobraćaja, metode detekcija anomalija pokazuju značajne prednosti u odnosu na tradicionalno korišćene metode inspekcije paketa zasnovane na potpisu (eng. signature-based detection). Razvijena metoda je zasnovana samo na osnovnom skupu atributa o mrežnim komunikacijama prikupljenih posredstvom NetFlow protokola, što je čini veoma pogodnom za praktičnu primenu u realnom vremenu. Iako ove osnovne informacije nisu dovoljne za primenu tehnika mašinskog učenja, njihovom agregacijom se dobijaju dodatni atributi koji su efikasni za primenu tehnika detekcije anomalija zasnovanih na entropiji. Dodatnom multivarijantnom analizom se osim detekcije anomalija omogućava i klasifikacija uočenih komunikacionih obrazaca. Uz prezentaciju najvažnijih elemenata primenjene metodologije, prikazana je modularna arhitektura realizovanog sistema i rezultati primene u realnom vremenu nad stvarnim saobraćajem u univerzitetskoj mreži.

Utorak, 27. jun 2023. u 10h, na platformi webex computing.math.rs/meet

Nevena Ćirić
Sveobuhvatni pregled i poređenje postojećih prediktora neuređenosti proteina sa analizom mogućnosti retreniranja modela

Apstrakt

Methodology, performance and retrainability survey of intrinsic disorder predictors

Intrinsically disordered proteins and regions are widely distributed within most proteomes. Recent studies show that they are associated with many essential biological processes and a broad range of human diseases. Given the prevalence of disordered proteins and the growing acknowledgement of their functional relevance, considerable effort has been made by the bioinformatics community to provide computational tools to predict protein disorder. To date, based on various characteristics of protein disorder, along with variety of diverse computational approaches, numerous disorder predictors have been developed.

Over the past decade several review papers examining intrinsic disorder predictors have been published. All these papers have played a significant role in stimulating and greatly facilitating the development of this actively growing field by pinpointing the potential room for improvement. Inspired by these, in this work we aim to integrate the relevant information regarding the existing intrinsic disorder predictors from the corresponding research papers in a novel review, including latest prediction tools. In addition, for each disorder predictor, we examined the possibility of their retraining using different datasets.

Here, we present an overview of 23 protein disorder prediction methods, including the thorough analysis of their advantages and weaknesses which derive from their different computational approaches. Regarding this, we precisely describe the methodology used for building the models and categorize them by different classification schemes. The performance of these models is presented by their scores from the most recent CAID competition. Additional contribution of this work is the models’ retraining availability analysis. We describe in detail the predictors’ implementation source code (if available) and propose a way around to overcome the obstacles with retraining procedure (if possible). This insight might be very useful, since older models were trained on significantly smaller datasets compared to the newer ones, due to the increase in the number of experimentally annotated disorder proteins with time. With respect to this, we discuss in detail the possibility of retraining the models on a different (bigger, novel) dataset in order to perform full-scale comparison of their expression power in delineating disorder in proteins.

Sreda, 31. maj 2023. u 18h, u učionici N153

Milena Šošić
Modelovanje moralnih i emocionalnih komponenti jezika u klasifikaciji konverzacionih tekstova
(predstavljanje teme doktorske disertacije)

Apstrakt

Конверзационе текстуалне поруке представљају један од главних начина комуницирања у савременом друштву. Са развојем информационих технологија у претходној деценији, појавили су се бројни алати за комуницирање као што су поруке електронске поште, поруке као објаве и коментари на друштвеним мрежама, поруке у алатима за инстант комуницирање и поруке за комуницирање са алатима за аутоматско генерисање одговора. Сви ови типови порука имају своју структуру и могу се класификовати, било појединачно или као скупови порука који припадају једном конверзационом низу, у различите врсте категорија. Oбласт рачунарске лингвистике је за неке светске језике већ постигла одређени напредак у откривању моралних и емоционалних компоненти у тексту, али су потребна додатна истраживања да би се у потпуности разумели ови сложени и апстрактни концепти, као и њихова међусобна повезаност. За српски језик овакви ресурси још увек нису развијени. Класификовање конверзација коришћењем карактеристика моралних и емоционалних компоненти језика представља задатак који још увек није довољно истраживан. Предмет овог истраживања биће анализа утицаја моралних и емоционалних компоненти језика у задацима класификације конверзационих текстова. Утицај ових компоненти ће се мерити по степену њиховог доприноса и значаја на успешност класификације односно на успешност поделе података у категорије дефинисане задатком класификације. Применом математички заснованих метода за проналажење значајних атрибута класификације утврдиће се у којој мери морални и емоционални атрибути узимају учешће у овако утврђеним скуповима. У оквиру класификације конверзационих текстова на енглеском језику биће предложен нови, хибридни приступ који користи методе машинског учења за класификацију ове врсте података. У оквиру истраживања биће развијени нови ресурси и предложене технике које би омогућиле примену предложеног приступа на конверзационе текстове написане на савременом српском језику.

Ponedeljak, 22. maj 2023. u 18h, u učionici 718

dr Josep Torrellas
Cloud Computer Architecture of the Next Decade

Apstrakt

The Cloud of the next decade will be defined by the need to process vast swaths of data for insights in a timely manner. Minimizing data movement to curtail energy consumption and increasing computation efficiency will be the overriding constraints. The compute infrastructure will be a seamless hierarchy of compute centers from edge to geo-distributed mega-datacenters. Each compute center will contain a large number of heterogeneous hardware accelerators, and tasks of unprecedentedly small granularity will ship computation to where data is. In this talk, I will describe some of the research on computer systems that my group is performing to attain this vision. We are developing programmable accelerators organized into ensembles, heterogenous memory systems, smart network interfaces and secure hardware structures. The combination of all these technique is likely to deliver large improvements in performance and energy efficiency.

Bio: Josep Torrellas is the Saburo Muroga Professor of Computer Science at the University of Illinois Urbana-Champaign (UIUC). He is the Director of the SRC/DARPA ACE Center for Evolvable Computing, past Co-Leader of an Intel Strategic Research Alliance (ISRA) on Computer Security, and past Director of the Illinois-Intel Parallelism Center (I2PC). His research interests are computer architectures for shared-memory multiprocessors and parallel computing. Some of his contributions include thread-level speculation (TLS) architectures, the Bulk Multiprocessor concept, deterministic record and replay mechanisms, process variation mitigation techniques, and hardware defenses against speculative execution attacks. In addition, he has contributed to several experimental multiprocessor designs such as IBM’s PERCS Multiprocessor, Intel’s Runnemede Extreme-Scale Multiprocessor, Illinois Cedar, and Stanford DASH. He has graduated 48 PhDs, who are now leaders in academia or industry.

Utorak, 16. maj 2023. u 12h, u učionici 718

dr Avi Mendelson
The impact of advanced machine learning algorithms on research and our society

Apstrakt

The use of advanced machine learning algorithms started having a tremendous impact on almost every aspect of our life, including: Politics, education, the way we conduct science, and more.

I will start my talk with a short description of some of the great opportunities the new technologies might open, with a discussion of potential dangers we may face. The rest of my talk will focus on what the academaia should do and what computer scientists can contribute to reducing the probability that the new technology will harm our society.

Biografija predavača: Avi Mendelson is a Professor at the CS department Technion and an IEEE Fellow. He has industrial and academic experience in various areas, such as Computer architecture, Hardware Security, Hardware accelerators, and Architecture for machine learning.

He graduated from the CS department, Technion, BSC (1979) and MSC (1982) and got his Ph.D. (1990) from the University of Massachusetts at Amherst (UMASS).

Among his industrial jobs, he worked for 11 years as a senior computer architect and principal engineer at Intel; he managed the academic-relations activities for Microsoft R&D Israel for four years, focusing on students’ innovation and research collaborations. In addition, he served as a PI in various EU projects, among them the EuroLab4HPC aiming at encouraging technology transfer between academia and startups.

Prof. Avi Mendelson has published more than 130 papers in refereed journals. And he holds more than 30 USA patents. He was a member of the ACM Europe council team and the IEEE Computer Society Board of Governors (BoG) and served as a second VP. Currently, manage the Computer Society Israeli Chapter in Israel.

Četvrtak, 6. april 2023. u 18h, na platformi webex computing.math.rs/meet

Jelena Marković
Kelerova hipoteza

Apstrakt

Kelerova hipoteza tvrdi da u bilo kom popločavanju n-dimenzionalnog Euklidskog prostora identičnim hiperkockama, postoje dve hiperkocke koje dele celo (n-1)-dimenzionalno lice. Od 1930. godine kada je Keler postavio hipotezu do sada, razvijena je matematička teorija koja nam omogućava proveru tačnosti hipoteze i koja definiše pojam Kelerovog grafa, a potom svodi tačnost hipoteze na postojanje klike veličine 2n u Kelerovim grafovima konstruisanim za dimenziju n. U radu koji je poslužio kao osnova za ovo predavanje akcenat je na dimenziji 7, jer je za dimenzije manje ili jednake od 6 pokazano da je hipoteza tačna. Posmatramo tri Kelerova grafa G7,3, G7,4 i G7,6. Hipoteza je netačna u ovoj dimenziji akko bar jedan od ovih grafova sadrži kliku veličine 27 = 128. Na predavanju će biti predstavljen automatski metod da se reši ova hipoteza kodiranjem postojanja klika iskaznom formulom. Dodatno, biće razmatrane tehnike razbijanja simetrija specifično konstruisane za ovaj problem, kako bi se smanjilo vreme rada SAT rešavača.

Četvrtak, 12. januar 2023. u 18h, na platformi webex computing.math.rs/meet i u učionici 718

Biljana Stojanović
Primena metoda istraživanja podataka na upotrebu kodona radi klasifikacije proteinskih sekvenci i određivanja međusobne sličnosti virusa
(predstavljanje teme doktorske disertacije)

Apstrakt

Одређивање склоности појединачних (група) организама (или њихових вируса) као и функционалних карактеристика протеина унутар организма према употреби кодона представља проблем који још увек није решен. Појам употребе кодона (енг. Codon Usage) се односи на склоност према одређеним синонимним кодонима аминокиселина унутар гена који кодирају протеине, скупа таквих гена у организмима, или врстама. Развијене су бројне биоинформатичке методе (мере) за анализу употребе кодона.

Основни циљ предложене теме је утврђивање утицаја (промене) коришћења кодона на класификацију протеинских (аминокиселинских) ниски, или делова ниски, на модел-систему појединачних вируса и/или група вируса, као и склоности вируса или њихових компоненти (гена) ка одређеним кодонима. Испитиваће се и постојањe мутација у сличним протеинима вируса и утицај тих мутација на евентуалну промену (функционалних) карактеристика протеина или самих вируса.

У раду ће бити анализирано више фамилија вируса са посебним освртом на фамилију коронавируса. Анализираће се постојеће методе за карактеризацију употребе кодона. Основне научне методе које ће бити коришћене у истраживању засноване су на моделима и алгоритмима у области истраживања података (одређивањa правила придруживања, класификацијe, кластеровањa, истраживањa образаца). Добијени рачунарски модели би требало да допринесу бољем разумевању карактеристика вируса, предвиђању сродних региона у вирусима, као и бољем резумевању утицаја мутација на нуклеотидном нивоу на промену карактеристика самог вируса.

Četvrtak, 15. septembar 2022. u 18h, na platformi webex computing.math.rs/meet

Lazar Vasović
Algoritmi konstrukcije sufiksnog niza

Apstrakt

Тражење једне ниске (шаблона) у другој (бази) основни је проблем који решавају многобројни алгоритми текста. Традиционални приступи, попут Бојер-Муровог и Кнут-Морис-Пратовог алгоритма, засновани су на припреми шаблона. Шаблони су, међутим, променљиви, док база углавном није, па су нови приступи засновани управо на њеној предобради. Припрема базе такође омогућава ефикасно решавање многих сродних задатака. Позната структура података која добро описује структуру базе јесте суфиксно стабло. Ипак, стабла су просторно захтевна и могу бити сложена. Уместо њих, ефикасније је чувати суфиксни низ, што је сортирани низ суфикса, односно низ индекса сортиран према лексикографском поретку суфикса који почињу тим индексима. Уз помоћни низ најдужих заједничких префикса, суфиксни низ ефикасно решава све проблеме претраге.

У раду је размотрен појам суфиксног низа, са посебном пажњом посвећеном његовој конструкцији, те су описана, имплементирана и упоређена четири алгоритма изградње. На првом месту је наивни алгоритам, који једноставно сортира индексе према лексикографском поретку одговарајућих суфикса, па је временске сложености O(n2 log n) за ниску дужине n. Следеће је дуплирање префикса (prefix doubling), које искључује многа поновљена поређења, чиме постиже сложеност реда O(n log n). Напослетку, реализовани су напреднији алгоритми DC3 (difference cover modulo 3) и SA–IS (suffix array–induced sorting), који су линеарне временске сложености, дакле O(n). Сви набројани приступи су упоређени како теоријски, тако и практично, табеларним и графичким приказом рада имплементација. Написана је и свита тестова, у циљу потврде коректности реализација.

Sve godine