서지주요정보
Efficient federated learning system for mobile artificial intelligence development = 효율적 모바일 인공지능 개발을 위한 연합학습 시스템
서명 / 저자 Efficient federated learning system for mobile artificial intelligence development = 효율적 모바일 인공지능 개발을 위한 연합학습 시스템 / Jaemin Shin.
발행사항 [대전 : 한국과학기술원, 2024].
Online Access 원문보기 원문인쇄

소장정보

등록번호

8043035

소장위치/청구기호

학술문화관(도서관)2층 학위논문

DCS 24012

휴대폰 전송

도서상태

이용가능(대출불가)

사유안내

반납예정일

리뷰정보

초록정보

The widespread adoption of mobile devices (e.g., smartphones, tablets) with Artificial Intelligence (AI) has led to breakthroughs in mobile services and applications. However, developing novel mobile AI often face challenges due to the difficulty of collecting real-world data, limited by factors such as budget constraints, privacy concerns, and small study groups. Federated Learning (FL) offers a solution to train AI models on large, private datasets from mobile devices, but it requires overcoming obstacles of slow learning and suboptimal model accuracy when trained on heterogeneous real-world devices. This thesis presents systematic approaches for 1) improving the time-to-accuracy and accuracy of FL on mobile devices and 2) providing guidance on selecting the most efficient FL algorithms to facilitate real- world deployment. Furthermore, this thesis extends the proposed techniques by applying them to the development of a novel mobile AI on user language-based mental health monitoring on smartphones, showcasing the practical implementation of the optimized FL methods on mobile scenarios.

최근 스마트폰, 태블릿 등 모바일 기기의 광범위한 보급과 인공지능(AI) 기술의 비약적인 발전은 인공지능 모바일 서비스의 발달로 이어져 인류의 일상 생활에 혁신을 가져왔다. 그러나 새로운 모바일 인공지능의 개발 과정에는 예산의 한계와 개인정보 보호 문제로 인공지능 학습에 필수적인 다양하고 많은 데이터를 수집하는 데 어려움이 있다. 이에 대한 해결 방안으로 대규모의 모바일 기기 위 데이터를 수집하지 않고 활용하여 인공지능 학습을 가능케 하는 연합학습(Federated Learning)이 제안되었지만, 이기종 및 데이터 이질성을 갖는 실제 모바일 기기 위 적용되었을 때 낮은 성능의 인공지능을 비효율적으로 학습하는 문제가 있었다. 이러한 문제의 해결을 위해, 이 논문은 1) 모바일 기기 위 연합학습의 학습 속도 및 결과 모델 정확도를 개선하고, 2) 가장 효율적인 연합학습 알고리즘을 상황에 따라 최적으로 선택하는 시스템을 제시한다. 또한, 이 논문은 제안된 시스템을 바탕으로, 스마트폰 위 사용자 언어 기반 정신 건강을 모니터링하는 새로운 모바일 인공지능을 효율적으로 개발하고 제안된 시스템의 활용성을 검증한다.

서지기타정보

서지기타정보
청구기호 {DCS 24012
형태사항 viii, 133 p. : 삽도 ; 30 cm
언어 영어
일반주기 저자명의 한글표기 : 신재민
지도교수의 영문표기 : Sung-Ju Lee
지도교수의 한글표기 : 이성주
Including appendix
학위논문 학위논문(박사) - 한국과학기술원 : 전산학부,
서지주기 References : p. 93-129
주제 Federated learning
Mobile computing
Systems for machine learning
Efficient machine learning
Applied machine learning
연합학습
모바일 컴퓨팅
기계학습 시스템
효율적 기계학습
기계학습 응용
QR CODE

책소개

전체보기

목차

전체보기

이 주제의 인기대출도서

MyDJ prototype.

System Overview of MyDJ.

FL on two datasets with different deadline configuration methods: (a) Temporalis contraction. (b) Mechanical waves propagation.

Raw data signals from two sensors when chewing. The bottom images in Figure 2.4a are from the video recorded during the experiment for demonstration. JPS, JES, JEE, and HMS stands for Jaw Protraction Start, Jaw Elevation Start, Jaw Elevation End, and Head Moving Start. Boxes indicate the distinctive signal patterns of each sensor that appears with chewing. The piezoelectric sensor shows low-freque

Time/frequency domain sensor responses at the sequence of different human activities.

Eating Episode Frames (EEFs) and eating episodes generation from three-second windows.

Example of camera usage for ground-truth label collection from the day-long study with various user activities; eating, working at a desk, conducting a chemical experiment, exercising at a gym, etc.

Specification of participants' eyeglass frame and head.

Illustration on the metrics of eyeglass frame and head of participants that are used in Table 2.1. Obs-to-obs is the straight-line distance between the left and right otobasion superius, which is the point of attachment of eyeglasses near the temporalis muscle MYK11].

Ground-truth label collection during the week-long study.

An example of eating episode coverage calculation. TP, TN, FP, and FN stand for True Positives True Negatives, False Positives, and False Negatives respectively.

Averaged results on a different number of selected features from the day-long study.

Episode-level F1-score per participant from the day-long study.

Per-participant results from the week-long study. All subgraphs share the same legend of Fig- ure 2.13b.

Power measurements of MyDJ on each functionalities.

User experience questionnaire scores on MyDJ and regular eyeglasses.

A comparison with previous eating detection studies. We compare following metrics: F1-score, Undetected Eating Episodes (UEE), False Alarms (FA), Power Draw (PD) in mW, Battery Capacity (BC) in mAh, and RunTime (RT) in hours.

Ordered gradient norm (GN) of samples from FL training rounds on two different datasets

FL on two datasets with different deadline configuration methods.

Overview of FedBalancer architecture and its operation in an FL round in seven steps.

Deadline efficiency (DDL-E) evaluation on different deadlines on two FL tasks.

Speedup and accuracy on five datasets with the real-world user data.

{w,lss,dss,p} parameters from the FedBalancer experiments in Table 3.1. Parameters with forward slashes indicate the different parameters from the multiple runs of experiments with three different random seeds.

Speedup and accuracy on different choice of FedBalancer parameters.

Performance evaluation of FedBalancer with different options on FEMI

Performance breakdown of FedBalancer into Sample Selection (SS) and Deadline Control (DC).

Collaboration of FedBalancer with three FL algorithms on FEMNIST dataset

Android devices for the testbed experiments.

Results from the testbed experiments.

On-device training latency (train) and communication latency (network) on different devices

Time and resource efficiency of syncFL (FedAvg [MMR+17a]) and asyncFL (FedBuff [NMZ+22]) algorithms on MNIST and Sent140 datasets. Env1 and Env2 indicate different client configurations consisting of IID / non-IID data split and different client training speed distribution. The results demonstrate the normalized performance of methods, with FedAvg+WFA's time and resource usage (RU) set as the base

Overview of Scout components and the flow of algorithm efficiency prediction.

Example source code of discrete event simulator for feature generation. FedAug+Cuersample and FedBuff algorithms are shown for syncFL and asyncFL, respectively. tpc,tcc,stcc indicate the generated features.

CDF ofSTCC feature generated on discrete event simulation of3,000 clients for FEMNIST dataset on four algorithms.

Example input for Scout's prediction model that involves client metadata and simulation-generated features.

Illustration of sequential client modeling procedure. Any RNNs or transformer architectures are applicable for the sequence encoder.

Experiment details on five datasets.

Time and Resource Usage (RU prediction results in Mean Absolute Error (MAE) on nive datasets

Design of Context-Aware Language Learning (CALL) methodology for FedTherapist. CALL maps the user's text to multiple temporal contexts (time, location, motion, and app). Each piece of text trains the relevant models on each context; for example, the text generated at daytime and home is used to train both the 'daytime' and 'home' models. CALL ensembles the context model outputs to determine the me

Mental health monitoring performance on different methods.

Performance of ensemble methods (EA: ensemble by averaging, EE: ensemble by weighted sum) of CALL and context models (TD: daytime, TN: nighttime, LH: home, Lo: other locations, Ms: stationary, MM: in motion, Ac: communication applications, and Ao: other applications) on four mental health monitoring tasks. All subgraphs share the same legend with Figure 5.2a and 5.2b.

FedBalancer on FedTherapist compared to the baseline approaches.

Screenshots of our data collection application. Participants completed three tasks to join the study. First, participants were asked to upload their voice sample (Task 1) to use VoiceFilter WMW+19] (Section 5.9.3) Participants then downloaded the model files used in our local data collection module (Task 2), and started the study (Task 3). We asked the participants for a daily mental health status

Word count of each participant's collected speech and keyboard input from the 10-day study.

Count of participant responses on different mental health status scores.

Loss and AUROC results on train, validation, and test samples over epochs from fine-tuning experiments in Section 5.3.1. Figure 5.6a and 5.6b indicate fine-tuning experiment of End-to-End BERT + MLP using a pre-trained RoBERTa-base [LOG+19] model with a learning rate of 0.0001. Figure 5.6c and 5.6d depict fine-tuning experiment of LLM on a pre-trained LLaMa-7B model with a learning rate of 0.0002.

Mental health monitoring performance on different text embedding methods at epoch 200.

Effect offine-tuning on text embedding methods on mental health monitoring performance, measured in AUROC on depression task.

Smartphone overhead on scenarios of FedTherapist: CPU denotes average per-core utilization post 5-min task execution. 1. restricted to measure individual Voice Activity Detector model execution from its API during duty cycling; 1. N/A measurements, as the model is not loaded on tested smartphones due to memory constraints.

Test loss on users' input word count from speech and keyboard experiments. The results show that FedTherapist is applicable on users with little speech and keyboard input (i 1,000 words on each input).