Multi-unmanned aerial vehicle track and intelligent reflecting surface shift joint optimization method and system
Abstract
The invention discloses a combined optimization method and system for the trajectory of multiple drones and the phase shift of an intelligent reflective surface, and establishes a wireless communication system model based on the assistance of multiple drones and the intelligent reflective surface. The smart reflector reflects to the base station, determines the channel model in the wireless communication system model and the energy consumption model of the UAV and the smart reflector, calculates the energy efficiency of the wireless communication system model; uses the K-mean clustering algorithm to cluster ground users , Using the priority experience playback MATD3 method to determine the position of the drone in each cluster, the drone and the intelligent reflecting surface assist the user communicating with the base station, the activated reflecting element of the intelligent reflecting surface and the phase of the activated reflecting element Shift, complete the joint optimization of the multi-UAV trajectory and the phase shift of the intelligent reflecting surface. The invention solves the problems of high communication delay and high power consumption in the existing offline optimization method.

Term
14.7 yearsto projected expiry
Projected expiry 25 May 2041, counted from filing; an application has no term until it is granted.
- Priority and filed
- Published
- Today
- Projected expiry
10 claims: 2 independent, 8 dependent
- 1L一种多无人机轨迹和智能反射面相移联合优化方法,其特征在于,包括以下步骤: S1、建立基于多无人机和智能反射面辅助的无线通信系统模型,用户发送的信号由安 装在无人机上的智能反射面反射到基站,确定无线通信系统模型中的信道模型以及无人机 和智能反射面的能量消耗模型,计算无线通信系统模型的能量效率; S2、基于步骤S1确定的信道模型以及无人机和智能反射面的能量消耗模型,利用K-均 值聚类算法将地面用户分簇,将能量效率作为优化目标,然后利用优先级经验回放MATD3方 法确定每个簇中无人机的位置,由无人机和智能反射面辅助与基站进行通信的用户,智能 反射面被激活的反射元件及被激活反射元件的相移,完成多无人机轨迹和智能反射面相移 的联合优化。
- 2根据权利要求1所述的方法,其特征在于,步骤S1中,基于多无人机和智能反射面辅 助的无线通信系统模型具体为:随机分布的用户数量为U,用户U被划分为K个区域,每个区 域内的用户数量为限,%+…+Uk+,・,+uK = U,智能反射面和无人机的数量为K,每个安装有智 能反射面的无人机服务一个区域中的用户;搭载在无人机上的智能反射面通过一个集成控 制器调整Μ个反射元件的相移;基站同时接收经过所有智能反射面反射的信号;基站的天线 数量为Ν,智能反射面的反射元件数为Μ,用户为单天线;基站的坐标为(xbs,y BS , ZbJ ,智能反 射面P的坐标为(X叫,,必啊,z吗),用户q的坐标为卜吗,歹啊/啊),一个区域中只有一个用户 发送信号,每一个用户发送的信号通过服务本区域的智能反射面反射到基站,并通过服务 其他区域的智能反射面反射到基站,同时参与通信的用户和智能反射面的数量都为K;智能 反射面的每一个反射元件独立调整入射信号的相移,同时保持幅度不变,智能反射面P的相 移矩阵是一个对角阵® p = diag (vj,对角线上的元素勺…煮外,…,屋“),(^表示 智能反射面P第m个反射元件的相移;智能反射面的被激活反射元件矩阵是一个对角阵Δ ρ = diag(uj ,对角线上的元素%= (δ。],…,6pm,・,6pM),总表示智能反射面P的第m个反射元 件是否激活。
- 3根据权利要求1所述的方法,其特征在于,步骤S1中,用户发送的信号通过无人机智 能反射面反射到基站分决策阶段、飞行阶段和信息传输阶段,决策阶段:无人机选择与哪个 用户进行通信,并选择进行信息传输的位置,智能反射面选择被激活的反射元件及其相移;飞行阶段:无人机以速度ν沿直线飞向在决策阶段中选择的信息传输位置;信息传输阶段: 无人机到达规定的位置之后悬停,在决策阶段中被选中的用户向智能反射面发送信号,智 能反射面的激活反射元件以对应的相位偏移将用户发送过来的信号反射到基站。
- 4根据权利要求1所述的方法,其特征在于,步骤S1中,将用户和智能反射面之间、智能 反射面和基站之间的信道建模为莱斯信道,从用户q到智能反射面P的信道Gpq为; 其中,Ρ表示参考距离d0=lm处的路径损耗,占是路径损耗指数,8是莱斯衰落因子,d1是 用户q和智能反射面P之间的欧几里得距离,Gpq是非视距传播分量,是阵列响应矢量,如 表示信号从用户q到智能反射面P的到达角的余弦值,人表示载波的波长,d表示天线间距; 从智能反射面P到基站的信道Fp为:其中,d2表示智能反射面p和基站之间的欧几里得距离,Fp是非视距传播分量, Fp,和%(%>3)是阵列响应矢量; 基站的接收信号y为: K K K K 丫小Σ贴十”“十Σ 2?:@心43十〃 q=l,q力jt p=l q=\ 7 q^-k p=l 其中,S为发送信号矩阵,Η为信道矩阵,hk是矩阵Η的第k列,Sk是矩阵S的第k行,η表示基 站端的加性白高斯噪声,方差为。2的循环对称复高斯变量; 将其他用户的干扰视为噪声,第k个用户的信干噪比SINRk为: SWR* =-----L-^i------------以Σ身小心户」+帆” 第k个用户的信息传输速率Rk为: %=log 1 K w.y F^&&G. 1t 乙1ρ Ρ Ρ Ρ» ____ 吗Σ £睁4G凶+|帆” 夕=1 其中,K为同时与基站进行通信的用户的数量,Wk为迫零检测滤波矩阵的第k行,gT为智 能反射面P与基站之间的信道矩阵的共辗转置,⑻p为智能反射面P的相移矩阵,*为智能反 射面P的被激活反射元件矩阵,Gpq为用户q和智能反射面p之间的信道,Gpk为用户k和智能反 射面P之间的信道,。2为噪声的方差。
- 5根据权利要求1所述的方法,其特征在于,步骤S1中,能量效率EE。为传输的数据量除 以无人机P和智能反射面P消耗的总能量,具体为: G EE =------p F +尸 其中,3强为无人机飞到指定位置消耗的能量,Gp为用户p经过无人机p和智能反射面p 辅助,向基站传输的数据量,£叫为智能反射面P消耗的能量,品明为无人机P的推进功率, T为无人机飞到指定位置需要的时间。
- 6根据权利要求1所述的方法,其特征在于,步骤S2中,使用K均值聚类算法对用户进行 分簇,具体为: 指定一个K值,从所有用户中随机抽取K个用户作为初始的聚类中心,然后计算其余的 所有用户与这K个初始聚类中心之间的距离,将距离聚类中心最近的用户划分到对应聚类 中心所属的簇,对于每个新形成的簇,聚类中心通过计算簇中用户的平均值得到,如果所有 簇的聚类中心与上一次计算得到的结果完全相同,聚类准则函数已收敛,所有用户划分到 正确的簇中。
- 7根据权利要求1所述的方法,其特征在于,步骤S2中,利用优先级经验回放MATD3方法 确定每个簇中无人机的位置,与基站进行通信的用户的位置,智能反射面被激活的反射元 件以及被激活元件的相移,完成多无人机轨迹和智能反射面相移的联合优化具体为: 将基于多无人机和智能反射面辅助的无线通信系统中无人机轨迹和智能反射面相移 的优化问题建模成一个马尔可夫博弈,每个安装有智能反射面的无人机作为一个智能体, 第k个智能体观测当前的环境状态基于策略网选择一个行为行为作用于环境后获得 奖励口,然后环境将以转移概率P (s'J Sq% ,…,a》转移到新的状态s';在每个时刻内,第k个智能体观测上一时刻无人机k的位置,以及第k个簇中与基站进行 通信的用户的位置作为状态训练策略网络的参数为。『将状态Sk作为输入,输出当前时 刻第k个无人机的位置,第k个簇中与基站进行通信的被激活用户向量,第k个智能反射面的 被激活元件向量以及相移向量作为行为ad第一训练价值网络和第二训练价值网络的参数 分别为3kl和3k2,两个训练价值网络将各个智能体观测到的联合状态S=(S],S2,・,Sk)和 采取的联合行为a=61/2,···,a/作为输入,分别输出联合状态-行为价值函数Qki(s,a「 a 2 , ··· ,a K , ω^)和(s ,a x ,a 2 , ··· ,a K , ω k2 ),目标策略网络将下一个状态s'作为输入,输出 下一个行为a、,用软更新的方式根据训练策略网络的参数9k更新目标策略网络的参数9'k, 第一目标价值网络和第二目标价值网络输入下一个状态-行为对(s' ,a'),分别输出 Q*j(s , q ,《;··,生和Q'k2(s' ,a'],a'2,,,a‘K, ω 用软更新的方式根据第一训练价 值网络的参数ω η和第二训练价值网络的参数ω女?更新第一目标价值网络的参数ω %和第 二目标价值网络的参数3^2; 将(s,a「a2,…,a『ri,r2,…作为智能体的一条经验存放在经验存储器中,当经 验存储器达到最大存储容量时,使用优先级经验回放的方法从中抽样小批量经验进行训 练,更新策略网络的参数和价值网络的参数。
- 8根据权利要求7所述的方法,其特征在于,每个无人机观测到的状态包括两个部分, 分别是上一时刻无人机k (k= {1,2,…,Κ})的位置 , G拼取 — {“切匕,必山匕’马以匕} ‘以及在第k簇 中,由第k个无人机和智能反射面辅助与基站进行通信的用户的位置, ={芯3%3%£},状态Sk的维度为六维;行为ak包括以下四个部分: i :当前时刻第k个无人机的位置=v k ,y^AV t ' z vav k }; ii:当前时刻在第k个簇中与基站进行通信的被激活用户向量Z; ={八虞,其中 的每一个元素表示相对应的用户是否激活,取值为0表示相对应的用户不激活,取值为1表 示激活,并且向量Z;中的元素应满足彳+或+…+品=1,表示在任一时亥。,一个簇中只有 一个被激活的用户; iii:当前时刻第k个智能反射面的被激活元件向量& ={凡号,…,α},其中的每一个 元素表示相对应的反射元件是否被激活,取值为。表示相对应的反射元件不激活,取值为1 表示激活,向量田中的元素应满足+6+…+端KM,表示每个智能反射面被激活的 元件数量应该在1〜Μ之间; iv:当前时刻智能反射面的相移向量式={见处…,氏},其中的每一个元素表示相对 应的反射元件的相移,1 <焉 < 乃; 4={α**, z;, &, } 奖励定义为能量效率EEk,q (s k ,a k ) =EE k o
- 9根据权利要求7所述的方法,其特征在于,使用策略梯度法更新第k个智能体的训练 策略网络的参数9 k为: ]F j ©)=万 X ▽4 &(s',w 4,…,<%)1d=w)▽ 4 徇⑸) 其中,J (%)是策略目标函数,F表示小批量抽样的大小,▽表示梯度算符,%是第k个智 能体学习到的策略,s/为利用优先级经验回放方法抽样的第j条经验中第k个智能体的状 态,片为第j条经验中第k个智能体的行为; 第k个智能体的训练价值网络1的参数3 1d 和训练价值网络2的参数3k2通过神经网络的 梯度反向传播来更新,损失函数分别为: ]F 履”1 =方 Σ 叼(丁 arg etQl -应(s', α:, Η,…,3 % )『 ]F 双2 =万X吗(丁argetQ; -Q k ^ J Μ:,必…,欧,你2)了 其中,Wj为重要性抽样权重,Targ eQ表示目标Q值; 目标策略网络的参数9'k,目标价值网络1的参数a %和目标价值网络2的参数a %分 别使用软更新的方式进行更新: 工一叫+口⑹叱 1d 3 k2+(「ak2 其中,a表示更新系数。
- 10一种多无人机轨迹和智能反射面相移联合优化系统,其特征在于,包括: 能量模块,建立基于多无人机和智能反射面辅助的无线通信系统模型,用户发送的信 号由安装在无人机上的智能反射面反射到基站,确定无线通信系统模型中的信道模型以及 无人机和智能反射面的能量消耗模型,计算无线通信系统模型的能量效率EE。; 优化模块,基于能量模块确定的信道模型以及无人机和智能反射面的能量消耗模型, 利用K-均值聚类算法将地面用户分簇,将能量效率作为优化目标,然后利用优先级经验回 放MATD3方法确定每个簇中无人机的位置,由无人机和智能反射面辅助与基站进行通信的 用户,智能反射面被激活的反射元件及被激活反射元件的相移,完成多无人机轨迹和智能 反射面相移的联合优化。
Independent claims10
246 paragraphs, as filed
A method and system for joint optimization method and system of multi-UAV trajectory and intelligent reflecting surface phase shift
[0001] The present invention belongs to the field of wireless communication technology, and specifically relates to a joint optimization method and system for multi-UAV trajectory and intelligent reflective surface phase shift.
Background technique
[0002] With the development of the Internet of Things technology, more and more devices need to access the communication network, and sometimes these devices will be distributed in a very large range, if only a single drone and a single intelligent reflective surface are the number The service provided by the huge communication equipment will undoubtedly bring a large communication load to the drone. In addition, the long-distance flight of the drone will consume a lot of time and energy, which will cause serious communication delays. The power consumption of drones poses challenges.
[0003] In order to improve low-latency and high-reliability services for user equipment, multiple drones and multiple intelligent reflecting surfaces can be used to assist communication, and the K-means clustering algorithm is used to divide ground users into several areas, Each drone equipped with an intelligent reflective surface serves users in a certain area. Under the premise of ensuring that good communication quality can be obtained, the trajectory of the drone and the intelligent reflective surface are jointly optimized through the use of multi-agent reinforcement learning algorithms. The phase shift maximizes the energy efficiency of the wireless communication system.
Summary of the invention
[0004] The technical problem to be solved by the present invention is to provide a joint optimization method and system for multi-drone trajectory and intelligent reflection surface phase shift in view of the above-mentioned deficiencies in the prior art, so as to solve the existing drone trajectory and intelligent reflection The off-line optimization method of surface phase shift has high communication delay and high power consumption.
[0005] The present invention adopts the following technical solutions:
[0006] A method for joint optimization of multi-UAV trajectory and intelligent reflecting surface phase shift includes the following steps:
[0007] S1. Establish a wireless communication system model based on the assistance of multiple drones and intelligent reflective surfaces. The signal sent by the user is reflected by the intelligent reflective surface installed on the drone to the base station, and the channel model in the wireless communication system model is determined. Energy consumption model of UAV and intelligent reflecting surface, calculating energy efficiency of wireless communication system model;
[0008] S2, based on the channel model determined in step S1 and the energy consumption model of the UAV and the intelligent reflecting surface, using
The K-means clustering algorithm clusters ground users, takes energy efficiency as the optimization goal, and then uses the priority experience playback MATD3 method to determine the location of the drones in each cluster, which is assisted by the drone and the intelligent reflective surface with the base station. For communication users, the phase shift of the activated reflective element and the activated reflective element of the smart reflective surface completes the joint optimization of the multi-UAV trajectory and the phase shift of the smart reflective surface.
[0009] Specifically, in step S1, the wireless communication system model based on the assistance of multiple drones and intelligent reflective surfaces is specifically: the number of randomly distributed users is U, and the user U is divided into K regions, and the number of users in each region The number of users is limited, %+...+Limit+...+Uk = U, the number of smart reflective surfaces and drones is K, and each drone equipped with a smart reflective surface serves users in an area; The intelligent reflecting surface on the drone adjusts the phase shift of the M reflecting elements through an integrated controller; the base station receives the signals reflected by all the intelligent reflecting surfaces at the same time; the number of antennas of the base station is N, and the number of reflecting elements on the intelligent reflecting surface is M , The user is a single antenna; the coordinates of the base station are (xbs, Ybs, Zbs), the coordinates of the smart reflective surface P are (seven, y% J, and the coordinates of user q are (%% %what/yang), a Only one user in the area sends a signal, each
The signal sent by each user is reflected to the base station through the intelligent reflecting surface serving this area, and reflected to the base station through the intelligent reflecting surface serving other areas. At the same time, the number of users participating in the communication and the number of intelligent reflecting surfaces are K; A reflective element independently adjusts the phase shift of the incident signal while keeping the amplitude unchanged. The phase shift matrix of the smart reflective surface P is a diagonal matrix. Nine = diag (vj, the element on the diagonal \/%), which means nine The phase shift of the m-th reflective element of the smart reflective surface; the activated reflective element matrix of the smart reflective surface is a diagonal matrix * = diag storage), the element on the diagonal% = (6 mountains..., 6pm,..., %?,<sup>8</sup>pm represents the<sup>111</sup>Whether each reflective element is activated.
[0010] Specifically, in step S1, the user sends a signal reflected by the reflecting surface to the base station UAV intelligent decision sub-stage, the flight phase and information transfer stage segment, the decision stage: UAV select which users to communicate with, and Select the location for information transmission, the intelligent reflective surface selects the activated reflective element and its phase shift; flight phase: the drone flies in a straight line at speed ν to the information transmission location selected in the decision-making phase; information transmission phase: unmanned After the machine arrives at the specified position, hovering, the selected user in the decision-making stage sends a signal to the intelligent reflecting surface, and the active reflecting element of the intelligent reflecting surface reflects the signal sent by the user to the base station with a corresponding phase shift.
[0011] Specifically, in step S1, the channel between the user and the intelligent reflecting surface, and between the intelligent reflecting surface and the base station is modeled as a Rice channel, and the channel Gpq from the user q to the intelligent reflecting surface P is:
<img file="CN113364495A_D0001.tif" />
[0013] where P represents the path loss at the reference distance d0=lm, & is the path loss index, 8 is the Rice fading factor, and Miao is the Euclidean distance between the user q and the smart reflecting surface ρ, or non The line-of-sight propagation component, the fog is the array response vector, % represents the cosine of the angle of arrival of the signal from the user q to the smart reflector ρ, human represents the wavelength of the carrier wave, and d represents the channel Fp of the antenna from the smart reflector P to the base station. : Spacing. [0014]
[0015]
[0016]
<img file="CN113364495A_D0002.tif" />
Among them, d2 represents the Euclidean distance between the intelligent reflecting surface ρ and the base station, the bean is the non-line-of-sight propagation component, and the sum is the array response vector;
[0017] The received signal y of the base station is:
[0018] Le=Wannian+ X old woman1+=2 Chen He ApG Wu s* + X £ Chens 4 households are called Sq+n
[0019] Among them, S is the transmission signal matrix, H is the channel matrix, which is the kth column of the matrix H, Sk is the kth row of the matrix S, n represents the additive white Gaussian noise at the base station, and the variance is. 2 cyclic symmetric complex Gaussian variables;
[0020] Regarding the interference of other users as noise, the signal-to-interference-to-noise ratio SINRk of the k-th user is: κ 2
[0021] SIN& --
[0022] The information transmission rate of the kth user is divided into:
<td rowspan="2">[0023] &=log 1+-X</td><td></td><td colspan="2"><sub>1</sub>P=1</td><td>2 、</td>
<td colspan="2"></td><td colspan="2">2+Fan Re/</td>
[0024] Among them, K is the number of users communicating with the base station at the same time, and sub-female is the k-th row of the zero-forcing detection filter matrix, is the co-transposition of the channel matrix between the smart reflector P and the base station, and is the smart reflector The phase shift matrix of plane P, Ap is the activated reflective element matrix of smart reflective plane P, Gpq is the channel between user q and smart reflective plane ρ, and Gpk is the channel between user k and smart reflective plane P. 2 is the variance of noise.
[0025] Specifically, in step S1, the energy efficiency EE. Divide the amount of transmitted data by the total energy consumed by the drone ρ and the intelligent reflecting surface ρ, which is specifically:
[0026] £=] Seven 5 Steps, Strict Seventeen Out 4
[0027] Among them, /Ming is the energy consumed by the drone flying to the designated location, Gp is the amount of data transmitted to the base station by the user ρ through the assistance of the drone ρ and the intelligent reflecting surface ρ, and b is the consumption of the intelligent reflecting surface P The energy is the propulsion power of the UAV P, and Τ is the time required for the UAV to fly to the designated position.
[0028] Specifically, in step S2, clustering the users using the K-means clustering algorithm is specifically as follows:
[0029] Specify a value of K, randomly select K users from all users as the initial cluster centers, and then calculate the distance between all the remaining users and the K initial cluster centers, and set the closest distance to the cluster centers. Users are divided into clusters to which the corresponding cluster center belongs. For each newly formed cluster, the cluster center is obtained by calculating the average value of the users in the cluster. If the cluster centers of all clusters are exactly the same as the result obtained in the previous calculation, the cluster The class criterion function has converged, and all users are divided into the correct clusters.
[0030] Specifically, in step S2, the priority experience playback MATD3 method is used to determine the location of the drone in each cluster, the location of the user communicating with the base station, the reflective element on which the smart reflective surface is activated, and the value of the activated element Phase shift, complete the joint optimization of multi-UAV trajectory and intelligent reflective surface phase shift specifically as follows:
[0031] The optimization problem of the UAV trajectory and the phase shift of the intelligent reflecting surface in the wireless communication system assisted by multiple UAVs and intelligent reflecting surfaces is modeled as a Markov game, and each unmanned intelligent reflecting surface is installed As an agent, the k-th agent observes the current state of the environment, Sy also chooses a behavior based on the strategy to act on the environment to obtain the reward, and then the environment will transfer with the transition probability P (s 11 Sq% ,...,a To new state s';
[0032] In each moment, the k-th agent observes the position of the drone k at the previous moment, and the position of the user communicating with the base station in the k-th cluster as the parameters of the state training strategy network. "Take the state Sk as input, output the position of the k-th UAV at the current moment, the vector of the activated user communicating with the base station in the k-th cluster, the vector of the activated element and the phase shift vector of the k-th smart reflector as The parameters of the first training value network and the second training value network of behavior ad are 3kl and 3k2, respectively. The two training value networks combine the joint states observed by each agent S=(S),S2,···, Sand The joint behavior taken a=...,a is used as input, and the joint state-behavior value function Qki (s, ··· ,a<sub>K</sub>, ω^) and (s,a<sub>x</sub>,a<sub>2</sub>, ···,A<sub>K</sub>, ω<sub>k2</sub>), the target strategy network takes the next state s as input, outputs the next behavior a, and uses soft update to update the parameters of the target strategy network according to the parameters of the training strategy network 9k
%', the first goal value network and the second goal value network input the next state-behavior pair (s',a'), respectively output Q% (s' ,aj ,a\,...,a[, 3 %) And Q'k2(s', aj, a\,...,a1, 3 ), use the soft update method to update the first target according to the parameter 3 kl of the first training value network and the parameter 3 k2 of the second training value network The parameter ω% of the value network and the parameters 3 and 2 of the second target value network;
[0033] Store..., a "ri, r2,... as an experience of the agent in the experience memory, when the experience memory reaches the maximum storage capacity, use the priority experience playback method to sample small batches of experience for training and update The parameters of the strategy network and the parameters of the value network.
[0034] Further, the state observed by each UAV includes two parts, which are the position of the UAV k (k = {1,2 ,...,K}) at the previous time, and &nine={dagger Nine, virtual 3 z win}, and in the k-th cluster, the position of the user who is assisted by the k-th UAV and the intelligent reflecting surface to communicate with the base station, = {¥£, sho produced vines}, the dimension of the state sk It is six dimensions; behavior includes the following four parts:
[0035] current time the first <location of the drone" = {Work "in,% also /%};
[0036] ii: the vector of activated users communicating with the base station in the k-th cluster at the current moment
Z; ..., Dance}, each element of which indicates whether the corresponding user is activated or not, the value is. Indicates that the corresponding user is not activated. A value of 1 indicates activation, and the vector Z; the elements in should satisfy ten...Ten dances=1, which means that at any one time, there is only one activated user in a cluster;
[0037] iii: The vector flute of the activated element of the k-th smart reflecting surface at the current moment = {4, Living,...", each element in it indicates whether the corresponding reflecting element is activated, and a value of 0 indicates phase The corresponding reflective element is not activated. A value of 1 means activation, and the elements in the vector & should satisfy 1 +6; +...+heartΚ must, which means that the number of activated elements for each smart reflective surface should be between 1 and Μ .
[0038] iv: The phase shift vector of the smart reflective surface at the current moment = {see here...here}, each element of which represents the phase shift of the corresponding reflective element, 1«%W is not;
[0039] &*={cwin
[0040] Reward is defined as energy efficiency EEk, rk(Sk, a/=ΕΕ[0041] Further, using the strategy gradient method to update the parameter 9k of the training strategy network of the k-th agent is:
[0042] QJ(4)=Stone Σ dagger must (/, music, ...&Bier%)Q% area)
[0043] Among them, J (%) is the strategy objective function, F is the size of the mini-batch sampling, and V is the gradient operator, which is the strategy learned by the k-th agent, and the bottom is the sampling using the priority experience replay method The state of the k-th agent in the j-th experience, and a is the behavior of the k-th agent in the j-th experience;
[0044] The parameter 3k1 of the training value network 1 and the parameter 3k2 of the training value network 2 of the k-th agent are updated through the gradient back propagation of the neural network, and the loss functions are:
]F
[0045] fos% = <sub>7T</sub>z Is it (T arg eQ-place (s,, * Cheng,..., exhibition)> one
[0046] loss^^-YwjiTargeig/-Beetle (W,α;,Take...,To@plantΝ
[0047] Wherein, Wj is the importance sampling weight, and D arg eg represents the target Q value;
[0048] The parameter O'of the target policy network, the parameter 3'of the target value network 1 and the parameter ω of the target value network 2 are respectively updated in a soft update manner:
[0049]. ', -α9, + ("α)",
Κ Κ Κ
[0050] ω<ί-a3ki+("a) ω
[0051] 3 wu...Limited + (1/) 3%
[0052] Wherein, a represents the update coefficient.
[0053] Another technical solution of the present invention is a joint optimization system for multi-UAV trajectory and intelligent reflecting surface phase shift, including:
[0054] The energy module establishes a wireless communication system model based on the assistance of multiple drones and intelligent reflective surfaces. The signal sent by the user is reflected by the intelligent reflective surface installed on the drone to the base station, and the channel model in the wireless communication system model is determined. As well as the energy consumption model of the UAV and the intelligent reflecting surface, the energy efficiency EE of the wireless communication system model is calculated. ;
[0055] The optimization module, based on the channel model determined by the energy module and the energy consumption model of the UAV and the intelligent reflecting surface, uses the K-means clustering algorithm to cluster the ground users, takes the energy efficiency as the optimization target, and then uses the priority The empirical playback MATD3 method determines the position of the drone in each cluster. The drone and the intelligent reflective surface assist the user communicating with the base station. The activated reflective element of the intelligent reflective surface and the phase shift of the activated reflective element are completed. Joint optimization of UAV trajectory and phase shift of intelligent reflecting surface.
[0056] Compared with the prior art, the present invention has at least the following beneficial effects:
[0057] A joint optimization method of multi-UAV trajectory and intelligent reflecting surface phase shift,
[0058] The establishment of the channel model and the energy consumption model is to calculate the energy efficiency, the maximization of energy efficiency is taken as the optimization goal to train the neural network, and finally the neural network can learn a strategy that allows the wireless communication system to obtain the maximum energy efficiency. Using the priority experience playback MATD3 method to optimize the UAV trajectory and the phase shift of the intelligent reflecting surface can make the UAV and the intelligent reflecting surface adaptively adjust their own strategies according to the changes of the environment, and has strong robustness.
[0059] Further, by dividing the user into multiple areas, each area is equipped with a UAV equipped with a smart reflective surface to provide services to the user, which can avoid the high power consumption of the UAV due to long-distance flight. The problem of high communication delay.
[0060] Further, in the decision-making stage, the drone in each area chooses which user to communicate with, and selects the location of information transmission. The intelligent reflective surface selection needs to activate the reflective element and determine the phase shift of the activated element. During the flight phase, the UAV flies in a straight line to the information transmission position determined in the decision-making phase. In the information transmission stage, the selected user in the decision-making stage sends a signal, and the intelligent reflecting surface reflects the signal sent by the user to the base station.
[0061] Further, establishing a suitable channel model is the basis for accurately calculating the information transmission rate, and the energy efficiency of the system can be further calculated after the information transmission rate is obtained.
[0062] Further, designing the trajectory of the UAV and the phase shift of the intelligent reflecting surface with energy efficiency as an optimization goal can achieve the goal of maximizing the energy efficiency of the system.
[0063] Further, when the UAV needs to serve a large area, in order to improve the communication quality and save the energy of the UAV, users need to be clustered, and each UAV installed with a smart reflector serves For users in a cluster, the drone flies within the coverage area of the cluster to provide services for users in the cluster.
[0064] Further, the UAV and the intelligent reflecting surface in each cluster are used as an agent, and each agent uses distributed execution of centralized training to learn, which can realize experience sharing and quickly learn to the highest energy efficiency of the system. Optimal strategy. Using the priority experience replay method to extract samples from the experience memory can learn more frequently from experiences with higher learning value and improve learning efficiency. The TD3 algorithm can solve the problem of overestimation of the Q value, so that the value network can make an accurate assessment of the value of the state "behavior pair."
[0065] Further, the channel state is related to the location of the user and the UAV, and the channel state is an important basis for determining the best position of the UAV for information transmission and the phase shift of the intelligent reflecting surface.<sub>k</sub>Set to the location of the drone at the last moment and the location of the user communicating with the base station, so that the agent can learn the implicit relationship between the location of the drone and the user's location and the channel state, so that the state can be directly changed.<sub>k</sub>Map to the behavior that maximizes energy efficiency a<sub>k</sub>, Without obtaining accurate channel state information. By taking the position of the drone, the activated element matrix of the smart reflective surface and the phase shift matrix as behavior a<sub>k</sub>, The intelligent reflecting surface can establish a high-quality line-of-sight propagation link between the user and the base station, and reflect the signal sent by the user to the base station.
[0066] Further, by obtaining the gradient of the strategy objective function and adjusting the parameters of the training strategy network to maximize the Q value, a strategy that can map the state to the optimal behavior can be found. Using the gradient descent method to update the parameters of the training value network to minimize the loss function, the value network can make an accurate assessment of the value of the state behavior pair. The parameters of the target strategy network and the target value network are updated in a soft update manner to improve the stability of the algorithm.
[0067] In summary, the present invention uses multiple drones and multiple intelligent reflective surfaces to assist in communication, and uses the K-means clustering algorithm to cluster users, and each drone and intelligent reflective surface serves one cluster Users in the middle; the priority experience playback MATD3 method can enable agents to learn strategies adopted by other agents through centralized training, and experience sharing among each agent, so as to quickly realize the correlation of multiple drone trajectories and intelligent reflective surfaces. The joint optimization of shifting maximizes the energy efficiency of the system.
[0068] The technical solutions of the present invention will be further described in detail below through the accompanying drawings and embodiments.
Description of the drawings
[0069] FIG. 1 is a system model diagram of the present invention;
[0070] FIG. 2 is a schematic diagram of a process of transmitting information from a user to a base station in the present invention;
[0071] FIG. 3 is a flowchart of K-means clustering algorithm;
[0072] FIG. 4 is a framework diagram of the MATD3 method of priority experience playback;
[0073] FIG. 5 is a schematic diagram of the influence of user transmit power on energy efficiency.
Detailed ways
[0074] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. . Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work shall fall within the protection scope of the present invention.
[0075] It should be understood that when used in this and appended claims, the terms "include" and "include" indicate
The existence of the described features, wholes, steps, operations, elements, and/or components does not exclude the existence or addition of one or more other features, wholes, steps, operations, elements, components, and/or collections thereof.
[0076] It should also be understood that the terms used in the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0077] It should also be further understood that the term "and/or" used in the present invention and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes These combinations.
[0078] Various structural schematic diagrams according to the disclosed embodiments of the present invention are shown in the accompanying drawings. These figures are not drawn to scale, some details are enlarged and some details may be omitted for clarity of presentation. The shapes of the various regions and layers shown in the figure and the relative sizes and positional relationships between them are only exemplary. In practice, there may be deviations due to manufacturing tolerances or technical limitations, and those skilled in the art may deviate according to actual conditions. Areas/layers with different shapes, sizes, and relative positions can be designed as needed.
[0079] The present invention provides a method for joint optimization of multi-UAV trajectory and intelligent reflecting surface phase shift. Firstly, a wireless communication system model based on the assistance of multi-UAV and intelligent reflecting surface is established, and secondly, the optimization problem of trajectory and phase shift is established. For the non-convexity of, the priority experience playback MATD3 method (Multi Agent Twin Delayed Deep Deterministic policy gradient, MATD3) is proposed, which realizes the joint optimization of multi-UAV trajectory and intelligent reflective surface phase shift.
[0080] A joint optimization method of multi-UAV trajectory and intelligent reflecting surface phase shift of the present invention includes the following steps:
[0081] S1. Establish a wireless communication system model based on the assistance of multiple drones and smart reflective surfaces, and then discuss the channels and the energy consumed by the drones and smart reflective surfaces respectively;
[0082] The communication model is shown in FIG. 1. Assuming that the number of users randomly distributed in a certain range is U, these users are divided into K areas, and the number of users in each area is Uk, U]+...+Uk+ ,,+uK = U. The number of smart reflective surfaces and drones are both K, and each drone equipped with a smart reflective surface serves users in an area. The intelligent reflective surface mounted on the drone adjusts the phase shift of the M reflective elements through an integrated controller. The base station receives all the signals reflected by the intelligent reflecting surface at the same time. Assuming that the number of antennas of the base station is N, the number of reflective elements of the smart reflector is M, and the user is a single antenna. Suppose that the coordinates of the base station are (XB s, y BS, z BS), the coordinates of the intelligent reflecting surface P are BU 5M ah/call), and the coordinates of user q are BU ah/tao, ζ%). At a certain moment, only one user in an area sends a signal. The signal sent by each user will not only be reflected to the base station through the intelligent reflecting surface serving the area, but also reflected to the base station through the intelligent reflecting surface serving other areas, and participate in the same time. The number of communicating users and smart reflecting surfaces is both κ. Each reflective element of the smart reflective surface can independently adjust the phase shift of the incident signal while keeping its amplitude unchanged. The phase shift matrix of the smart reflective surface P is a diagonal matrix ® p = diag (Vp), diagonal The upper element 4 E bedroom":
[0083] (1 )
[0084] Among them, 9Pm represents the phase shift of the mth reflective element of the smart reflective surface P.
[0085] The activated reflective element matrix of the smart reflective surface is also a diagonal matrix A p = diag (uj, the element %6 on the diagonal must be:
[0086] %=(δρ],...small,...4)
[0087] Among them, 6pm indicates whether the mth reflective element of the smart reflective surface p is activated,
<sub>r Ί</sub> , il, the first reflective element of the smart reflective surface P is activated
[00881 δ =< (3) "[θ, the additional reflective element of the smart reflective surface ρ is not activated
[0089] Please refer to FIG. 2, the process of the user transmitting information to the base station is divided into three stages, specifically:
[0090] 1) Decision-making stage: The drone chooses which user to communicate with, and selects the location for information transmission, and the intelligent reflective surface selects the activated reflective element and its phase shift.
[0091] 2) Flight phase: The drone flies along a straight line at a speed ν to the information transmission position selected in the decision phase.
[0092] 3) Information transmission stage: After the drone arrives at the specified position, hovering at that position, the selected user in the decision-making stage sends a signal to the intelligent reflecting surface, and the active reflecting element of the intelligent reflecting surface is shifted in a certain phase. The mobile reflects the signal sent by the user to the base station.
[0093] The channel between the user and the intelligent reflecting surface and the intelligent reflecting surface and the base station is modeled as a Rice channel, and the channel from the user q to the intelligent reflecting surface P is 6 min", ρ, 9 = {1 ,...&...,Κ}, specifically:
[0094] G". Heel (4)
[0095] Among them, P represents the path loss at the reference distance d0=lm, & is the path loss index, 8 is the Rice fading factor, 4 = J(x dense-X called>+(a few 4 -%sj+(z -2%) 2 is the Euclidean distance between the user q and the intelligent reflecting surface p, and the acid G is the non-line-of-sight propagation component, each element of which is modeled as a cyclic symmetric complex Gaussian variable with zero mean and unit variance , Dew= is the array response vector, which represents the cosine value of the angle of arrival of the signal from the user q to the intelligent reflecting surface P.
The channel from the smart reflector P to the base station is ©", specifically:
[0096]
[0097]
[0098]
<img file="CN113364495A_D0003.tif" />
(6) (7) Among them, 2 = J(*g, I Xbs )2 +(Ming-Ybs)? + called -z§s, which represents the Euclid between the intelligent reflecting surface p and the base station The distance, the slice is a non-line-of-sight propagation component, each element of which is modeled as a cyclic symmetric complex Gaussian variable with zero mean and unit variance, and the month is the array response vector, specifically: [0099] F<sub>p</sub> =α<sub>Μ</sub>(φ<sub>Ρ</sub>2^<sub>Ν</sub>(φ<sub>ρ</sub>^
[0100] 4/%2) and inquiry (0") can be expressed as:
[0101] 4must@>2)=7male,...,e'
[0102] %)=1,e Ramie...J Fu"
[0103] Among them, P·2 and Pa represent the cosine value of the departure angle and the arrival angle of the signal, respectively.
[0104] Assuming that the transmission signal matrix is S, the channel matrix is H, and the transmission signal of user k (k={1,...,k,...,Κ}) is the received signal of the base station:
Κ £ Κ K
[0105] y^h<sub>k</sub>s<sub>k</sub>+LiTie+=Zhichen®AGp*10£Lihu two packs eight households 4%10 (9) p=lg=l, qwk p=]
[0106] Among them, is the kth column of the matrix H, and is the kth row of the matrix S, and n represents the additive white Gaussian noise at the base station side, with the mean value being 0 and the variance being. The cyclic symmetric complex Gaussian variable of 2 is "~CN(0q2).
[0107] In an uplink multi-user communication system, since multiple users transmit signals on the same frequency band at the same time, there is co-channel interference. In order to suppress the co-channel interference between users and successfully detect the signals sent by each user, the base station can use the zero-forcing detection algorithm. The purpose is to eliminate the interference between the signals transmitted by different antennas at the signal receiving end through linear transformation.
[0108] In order to recover at the base station side and eliminate the interference of signals sent by other users, the matrix Wzf is used as the inner product of the received signal y to obtain an equalized signal, namely
[0109]
[0110]
W<sub>ZF</sub>y=W<sub>ZF</sub>HS+W<sub>ZF</sub>n As the kth row of the matrix Wzf, the Asian woman should meet the following conditions: (10)
[0111]
[0112]
[0113]
[0114] The matrix Wzf should satisfy WzfH=1 as a unit matrix, specifically:
W<sub>ZF</sub>= (H<sup>h</sup>H) <sup>-1</sup>H<sup>h</sup> (11) (12) Assuming that the H column of the channel matrix is full rank, the estimated value jP of the transmitted signal at this time can be expressed as:
[0115]
[0116] <sup>l</sup>H<sup>H</sup>(lIS + n)^s + W<sub>ZF</sub>n (13) The estimated value of the transmitted signal obtained after the zero-forcing detector completely eliminates the interference between the transmitted signals of different users.
[0117] Regarding the interference of other users as noise, the signal-to-interference-to-noise ratio of the k-th user is: ,, 2
[0118]
SiNR» = (14) suction Σ p=\
[0119] The information transmission rate of the kth user is:
[0120] / = log is called: £use ®a for use __g Ming Σ project through the households to trap Pian Lei (15)
[0121] Among them, K is the number of users communicating with the base station at the same time, Wk is the kth row of the zero-forcing detection filter matrix, gr is the co-transposition of the channel matrix between the intelligent reflecting surface p and the base station, and ρ is the intelligent The phase shift matrix of the reflecting surface ρ, δ<sub>ρ </sub>Is the activated reflective element matrix of the smart reflective surface P, Gpq is the channel between the user q and the smart reflective surface ρ, and Gpk is the user k
The channel between CN 113364495 A and the intelligent reflecting surface P. 2 is the variance of noise.
[0122] The energy consumption in the multi-user uplink transmission system assisted by the drone and the smart reflector includes two parts, namely the energy consumed by the drone in flight and the energy consumed by the activated reflective element of the smart reflector. The propulsion power of the P-th UAV is: ί ί I1 Falling 20 ¥ 1
[. Net product L brother] + elimination + piece Jl +/- ten thousand; (
I (Υ "mouth£/ J /
[0124]Brother=§*/«3 Ο
[0125] corpse_ 0 + E) inch <sup>1</sup> 4ΪκΑ (17) (18)
[0126] where is the speed of the tip of the UAV rotor blade, ν° is the average induced speed of the rotor during hovering, x is the airframe drag ratio, κ is the air density, u is the solidity of the rotor, and Α is the rotor disc Area, 3 is the section drag coefficient, Q is the blade angular velocity, Τ is the rotor radius, "is the incremental correlation coefficient of the induced power, W is the weight of the drone. ν. is the speed of the ρth drone, the calculation process as follows:
[0127] (t-1) The position of the drone ρ at time (t-1) is Bo Ying, the position at the time of Ying/Kangbu is value // won), and the UAV P flies from the position at time (t-1) to The distance experienced by the position at time t is:
[0128] d<sub>p</sub>T9)
[0129] Assuming that the time elapsed by the drone flight is T, the speed ν of the p-th drone. Is: d
[0130] Dagger, two inches (20)
[0131] The energy consumed by the drone ρ flying to the designated location is:
[0132](21)
[0133] Let δ represent whether the m-th reflective element of the smart reflective surface p is activated, and P]rs represents the power consumed by each reflective element of pin ± IVO, then the power consumed by the entire smart reflective surface P is:
[0134] P3P=Ej 3 deepP twist 5(22)
Mi -C
[0135] The duration of the information transmission phase is τ, and the energy consumed by the smart reflective surface p during this time is:
[0136] «%⑵)
[0137] In the information transmission stage, in the k-th cluster, the user ρ is assisted by the drone ρ and the intelligent reflecting surface ρ, and the amount of data transmitted to the base station is:
[0138] G =R τ (24)
Ρ Ρ
[0139] Energy efficiency is the amount of data transmitted divided by the total energy consumed by the drone ρ and the intelligent reflecting surface ρ:
G<sub>a</sub>
[0140] EE<sub>p</sub> = - (25)% A Seventeen's ρ
[0141] S2, based on the channel model and energy consumption model in step S1, use the K-means clustering algorithm to cluster ground users, and then use the priority experience playback MATD3 method to determine the location of the drone in each cluster, by The UAV and the intelligent reflecting surface assist users who communicate with the base station. The activated reflecting element and the phase shift of the intelligent reflecting surface during the information transmission stage complete the joint optimization of the UAV trajectory and the phase shift of the intelligent reflecting surface.
[0142] Please refer to FIG. 3, the basic idea of the K-means clustering algorithm is to first specify a K value, randomly select K users from all users as the initial clustering center, and then calculate the relationship between all the remaining users and the K value. The distance between the initial cluster centers, and which cluster center the user is closest to, is divided into the cluster to which the cluster center belongs. For each newly formed cluster, its cluster center is obtained by calculating the average value of the samples in the cluster. If the cluster center of all clusters is exactly the same as the result obtained in the previous calculation, it means that the clustering criterion function has converged and all users are Are divided into the correct clusters.
[0143] After using the K-means clustering algorithm to divide the ground users into multiple clusters, a drone with a smart reflector can be placed in each cluster, and the drone will fly within the coverage of the cluster. , To provide services for users in the cluster. The priority experience playback MATD3 algorithm is used to jointly optimize the trajectory of multiple drones and the phase shift of the intelligent reflecting surface to maximize the energy efficiency of the system. The algorithm framework is shown in Figure 4. The optimization problem of the UAV trajectory and the phase shift of the smart reflection surface in the wireless communication system based on the multi-UAV and the intelligent reflection surface is modeled as a Markov game, and each UAV equipped with the smart reflection surface acts as a Markov game. The agent, the k-th agent observes the current state of the environment based on the strategy and also chooses a behavior to act on the environment to be rewarded, and then the environment will transfer to the new one with transition probability P (s \ | Sk, a),... State s'[0144] The state observed by each UAV consists of two parts, which are the position of the UAV k (k= {1, 2,...,Κ}) at the previous moment. win=5win, Yiying Tokang}, and in the k-th cluster, the position of the user who is assisted by the k-th UAV and the intelligent reflector to communicate with the base station, & £ = {dagger 3child3z £}, the dimension of the state Sk For six dimensions:
[0145] s* = {Cwin,} (26)
[0146] The behavior of the k-th agent is a vector with a dimension of (3+%+2><M), not the number of users in the k-th cluster, and the behavior includes the following four parts:
[0147] i: The position of the k-th UAV at the current moment. Win={x Win, Win};
[0148] ii: the vector Z of activated users communicating with the base station in the k-th cluster at the current moment; = {fee,;,...,? ;}, each element in it indicates whether the corresponding user is activated, a value of 0 indicates that the corresponding user is not activated, a value of ί indicates activation, and the elements in the vector Z should meet + 7; +...+ tired =1, it means that there is only one activated user in a cluster at any one time;
[0149] iii: The activated element vector of the k-th smart reflecting surface at the current moment, each element of which indicates whether the corresponding reflecting element is activated, a value of 0 indicates that the corresponding reflecting element is not activated, and the value is I means activation, and the elements in the vector should satisfy seed+%+...+^W Μ, which means that the number of activated elements for each smart reflective surface should be between 1 and Μ.
[0150] iv: The phase shift vector ratio of the smart reflective surface at the current moment = {discuss, when..., see}, where each element represents the phase shift of the corresponding reflective element, 1 w% w not.
[0151] (27)
[0152] Reward Port (s<sub>k</sub>,a<sub>k</sub>) Is defined as the energy efficiency EEy calculated by equation (25).
[0153] For a multi-agent system, each agent has six neural networks, namely training strategy network, target strategy network, first training value network, second training value network, first goal value network and second goal Value network.
At each time, the k-th agent observes the position of the drone k at the previous time, and the position of the user communicating with the base station in the k-th cluster as the parameters of the state Sy training strategy network. "Take the state Sk as the input, output the position of the k-th UAV at the current moment, the vector of the activated user communicating with the base station in the k-th cluster, the vector of the activated element and the phase shift vector of the k-th smart reflector as Behavior ad The parameters of the first training value network and the second training value network are 3 and 3, respectively. These two networks combine the joint state s=(S),S2,,Sk) observed by each agent and the joint behavior taken a= (aa...,a) as input, respectively output the joint state-behavior value function Q (s,aa...,a, η κι.1. Bη<sup>w</sup><sub>kl</sub>) And Qk2(s,a] #2,...,a3k2), the target policy network takes the next state s'k as input, outputs the next behavior a'k, and uses soft update to train the policy network parameters 9k update the parameter %'of the target policy network, the first target value network and the second target value network input the next state-behavior pair (s',a'), and output Q'ki (s',a',a' respectively) 2,···#'κ, 3%) and Q'k2(s',a',a'2,,a'K,3,k2), using soft update method according to the first training value network Parameter ω "and the parameter ω of the second training value network? Update the parameter ω% of the first target value network and the parameter of the second target value network 3
[0154] Store (s, a1, a?,..., q, j,..., s') as an experience of the agent in the experience memory, when the experience memory reaches the maximum storage capacity, use the priority experience playback The method is to sample small batches of experience for training, and update the parameters of the strategy network and the value network.
[0155] The probability of experience j being sampled is:
D<sup>7</sup> (28)
[0156]
Σ0
7=1
[0157] Among them, Y represents the importance of the priority, F represents the number of experiences extracted in small batches, Dj=1/rank(j), and rank(j) is the ranking of the j-th empirical learning value.
[0158] Importance sampling weights are:
[0160] E is the number of experiences stored in the experience memory, and ξ is the sampling weight coefficient.
[0161] Use the strategy gradient method to update the parameters of the training strategy network of the k-th agent. Female:
[0162] %Biyou)=:Strength 0Md (v,...&-, 1*)l+"Point Q%/) (30)
[0163] Among them, J (%) is the strategy objective function, represents the gradient operator, and prison is the strategy learned by the k-th agent, and Ou is the kth of the jth experience sampled by the priority experience playback method The state of the agent, & is the behavior of the k-th agent in the j-th experience.
[0164] The parameter 3k1 of the training value network 1 and the parameter 3k2 of the training value network 2 of the k-th agent are updated through the gradient back propagation of the neural network, and the loss functions are:
]F
[0165] /0¾. - Σ<sup>ar</sup>§-Qk^<sup>J</sup>ai,(31) Primary Six 1
[0166] ]F ,. % = Square number (arg eQ-&? (s,, ,must,...,service,ming) production (32)
[0167] Among them, Wj is the importance sampling weight, and arg represents the target Q value.
[0168] Ding argef& = /'+ η min@| (s,,ά;, af,..., Lv, Know), 22 (s', 0Eightα:,...,%; <sup>ω</sup>^)) (33)
[0169] The loss function represents the gap between the Q value output by the training value network and the target Q value. Using the gradient descent method to update the parameters of the training value network to minimize the loss function, the Q value output by the training value network will be very close to the target Q Value, so that the training value network can make an accurate assessment of the value of the state's behavior pair.
[0170] The parameter Ο'of the target strategy network, the parameter 3% of the target value network 1 and the parameter ω% of the target value network 2 are respectively updated in a soft update manner:
[0171] 97-α%+("α) O'(34)
[0172] 3 %-a 3<sub>1d</sub>+ (1-α) 31ι(35)
[0173] 3 %-α 3k2+("α) 3 wu(36)
[0174] Wherein, α represents an update coefficient.
[0175] In still another embodiment of the present invention, a method and system system for joint optimization of multi-drone trajectories and intelligent reflecting surface phase shifts are provided, and the system can be used to realize the above-mentioned joint optimization of multi-drone trajectories and intelligent reflecting surface phase shifts. Method and system method. Specifically, the method and system system for joint optimization of multi-UAV trajectory and intelligent reflecting surface phase shift include an energy module and an optimization module.
[0176] Among them, the energy module establishes a wireless communication system model based on the assistance of multiple drones and intelligent reflective surfaces. The signal sent by the user is reflected by the intelligent reflective surface installed on the drone to the base station to determine the wireless communication system model. The channel model and the energy consumption model of the UAV and the intelligent reflector are used to calculate the energy efficiency EE of the wireless communication system model. [0177] The optimization module, based on the channel model determined by the energy module and the energy consumption model of the UAV and the intelligent reflector, uses the K-means clustering algorithm to cluster the ground users, and then uses the priority experience to replay the MATD3 method to determine each The position of the drones in each cluster is assisted by the drone and the intelligent reflective surface to communicate with the base station. The activated reflective element of the intelligent reflective surface and the phase shift of the activated reflective element complete the trajectory and intelligence of multiple drones. Joint optimization of the phase shift of the reflecting surface.
[0178] In yet another embodiment of the present invention, a terminal device is provided, the terminal device includes a processor and a memory, the memory is used to store a computer program, the computer program includes program instructions, the processor is used to execute Program instructions stored in the computer storage medium. The processor may be a central processing unit (Central Processing Unit, CPU), or other general-purpose processors, digital signal processors (Digital Signal Processor, DSP), application specific integrated circuits (Application Specific Integrated Circuit, ASIC), ready-made programmable Field-Programmable GateArray (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc., which are the computing core and control core of the terminal, are suitable for implementing one or more instructions. It is suitable for loading and executing one or more instructions to realize the corresponding method flow or corresponding function; the processor described in the embodiment of the present invention can be used for the operation of the joint optimization method and system of multi-UAV trajectory and intelligent reflecting surface phase shift, including :
[0179] Establish a wireless communication system model based on the assistance of multiple drones and intelligent reflective surfaces. The signal sent by the user is reflected by the intelligent reflective surface installed on the drone to the base station, and the channel model and unmanned communication system model in the wireless communication system model are determined. machine
Calculate the energy efficiency of the wireless communication system model with the energy consumption model of the intelligent reflecting surface; based on the determined channel model and the energy consumption model of the UAV and the intelligent reflecting surface, the ground users are clustered using the K-means clustering algorithm, and the Energy efficiency is taken as the optimization goal, and then the priority experience is used to replay MATD3 to determine the position of the drone in each cluster. The drone and the intelligent reflective surface assist the user communicating with the base station, the reflective element on which the intelligent reflective surface is activated, and the Activate the phase shift of the reflective element to complete the joint optimization of the multi-UAV trajectory and the phase shift of the intelligent reflective surface.
[0180] In yet another embodiment of the present invention, the present invention also provides a storage medium, specifically a computer-readable storage medium (Memory), and the computer-readable storage medium is a memory device in a terminal device for storing Procedures and data. It can be understood that the computer-readable storage medium herein may include a built-in storage medium in the terminal device, and of course, may also include an extended storage medium supported by the terminal device. The computer-readable storage medium provides storage space, and the storage space stores the operating system of the terminal. In addition, the storage space also stores one or more instructions suitable for being loaded and executed by the processor, and these instructions may be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here may be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0181] One or more instructions stored in a computer-readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the method and system for joint optimization of multiple drone trajectories and intelligent reflector phase shifts in the above-mentioned embodiments; computer One or more instructions in the readable storage medium are loaded by the processor and execute the following steps:
[0182] Establish a wireless communication system model based on the assistance of multiple drones and intelligent reflecting surfaces. The signals sent by users are reflected by the intelligent reflecting surface installed on the drones to the base station, and the channel model and unmanned communication system models in the wireless communication system model are determined. The energy consumption model of the aircraft and the intelligent reflecting surface is used to calculate the energy efficiency of the wireless communication system model; based on the determined channel model and the energy consumption model of the UAV and the intelligent reflecting surface, the ground users are clustered using the K-means clustering algorithm, Taking energy efficiency as the optimization goal, and then using priority experience to replay MATD3 to determine the position of the drone in each cluster, the drone and the intelligent reflective surface assist the user communicating with the base station, the reflective element on which the intelligent reflective surface is activated, and The phase shift of the activated reflective element completes the joint optimization of the multi-UAV trajectory and the phase shift of the intelligent reflective surface.
[0183] In order to make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be described clearly and completely in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the description The embodiments are a part of the embodiments of the present invention, but not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings herein may be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work shall fall within the protection scope of the present invention.
[0184] The joint optimization algorithm of multi-UAV trajectory and intelligent reflection surface phase shift based on priority experience playback of MATD3 is summarized as follows: Algorithm: Based on priority experience, the multi-UAV trajectory and fault of MATD3 are simultaneously released.<sup>1</sup>Reflective surface phase shift is excellent
[0185] Initialization:
[0186] Randomly initialize the parameters of each agent training strategy network %&···,%, the parameters of the training value network, such as, drink, electricity?,··.,%, know, assign the parameters of the training network to The target network initializes the update frequency C of the target network parameters, the maximum number of time slots in an experiment? And e, the maximum number of experiments is: the size of the experience memory and the size of the small batch sampling to produce online learning: for e=l, 2, .. ., Epi do initialize S to the first state of the current state sequence fbr t=l, 2, ..., Time do For each agent, choose a behavior 4 according to state 5, get rewards , and then go to the next State number, store the area as an experience in the experience memory "1, where $ = area, town, α = (α<sub>ι</sub>,α<sub>2</sub>,-··,Ά<sub>κ</sub>), R = [i\,r<sub>2</sub>,--,r<sub>K</sub>), S -(5[,S2,...,sQ for k=l, 2,..., K do a) Use the sampling method of priority experience playback to sample small batches of samples from the experience memory, the number is F, hit, add , Yin, S1,/ = 1,2,-.,/<sup>7</sup> b) Calculate the target Q value,
Target. / =d+/7 hernia 11 (α servant/, male; g',...,α], or,...,***/
C) Loss function, ". It is called =! Z force (Ding arg history 2-& 3,, W,..., such as)> ,? (Ding argeQ-(bar, Yibi,%)) 2, through the gradient of the neural network J=l Backpropagation to update the parameters of the first training value network and the parameters of the second training value network
<td>[0187]</td><td>d) According to the strategy gradient Q4&M, Yeah...&, %chuan" from nine%) Update the parameters of the training strategy network ae) If Time%c = 1, then use the soft update method to update the parameters of the target strategy network 4 , The parameter %I of the first target value network, or the parameter of the second target value network?, 4 +(1-α)4, just ί-α such as +(1-a) or, D sets the next state Current state S -s'end forend forend for</td>
<td>[0188]</td><td>The simulation parameters are set as follows:</td>
<td>[0189]</td><td>The parameter value is the number of users N 60 The side length of the area where the users are distributed. 500m Each smart reflective surface contains the number of reflective elements M 100 The path loss at a reference distance of 1 meter 2 Path loss index and 2 Rice fading factor 10dB Carrier wavelength% 0,1667 Antenna spacing d 0,0833</td>
<td>User's transmit power 4 plus</td><td>1W</td>
<td>The speed of the tip of the UAV rotor blade ί;""</td><td>120m/s</td>
<td>The average induced speed of the rotor during hover</td><td>4.03m/s</td>
<td>Airframe resistance ratio</td><td>0.6</td>
<td>Air density K</td><td>1.225kg/m<sup>3</sup></td>
<td>Density turned"</td><td>0.05</td>
<td>/Rotor disc area</td><td>0.503m<sup>2</sup></td>
<td>Section resistance coefficient 4</td><td>0.012</td>
<td>Blade angular velocity C</td><td>300m/s</td>
<td>[0190] Rotor radius Y</td><td>0.4m</td>
<td>Induced power increment correlation coefficient A</td><td>0.1</td>
<td>1E for drones</td><td>2kg</td>
<td>Discount factor? 7</td><td>0.9</td>
<td>Learning rate</td><td>0.001</td>
<td>Neural network parameter update coefficient a</td><td>0.1</td>
<td>The number of small batches of experience κ</td><td>64</td>
<td>Maximum number of experiments</td><td>100</td>
<td>The maximum number of moments in an experiment</td><td>2000</td>
<td>The size of experience memory Ε</td><td>1000</td>
[0191] Please refer to FIG. 5, which uses the MADDPG method, the MATD3 method, and the priority experience replay MADDPG method. When the priority experience replays the MATD3 method, the energy efficiency of the system changes with the user's transmission power. It can be seen from the figure that the energy efficiency of the system is higher when the priority experience playback method is used than when it is not used, and the energy efficiency of the system when the MATD3 method is used is higher than when the MADDPG method is used. This is because the experience memory is used when the priority experience playback is used. Of higher learning value
The probability of experience being sampled will increase. Learning from these experiences will improve learning efficiency. The MATD3 method can overcome the problem of overestimation of the Q value and enable the value network to accurately assess the value of the state "behavior pair. In addition, when the user When the transmission power of the system increases, the amount of data transmitted will increase, so the energy efficiency of the system will increase.
[0192] In summary, a method and system for joint optimization of multi-drone trajectory and intelligent reflector phase shift of the present invention takes into account the uplink wireless communication system based on multi-drone and intelligent reflector assistance. The ground users are divided into clusters, each cluster is assigned a drone with a smart reflective surface to provide services to the users in the cluster, and then the priority experience playback MATD3 method is used to complete the drone trajectory and the smart reflective surface in each cluster. The joint optimization of shifting in order to achieve the maximum energy efficiency of the system.
[0193] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, this application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, this application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0194] This application is described with reference to flowcharts and/or block diagrams of methods, devices (systems), and computer program products according to embodiments of this application. It should be understood that each process and/or block in the flowchart and/or block diagram, and the combination of processes and/or blocks in the flowchart and/or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing equipment to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing equipment are used to generate It is a device that realizes the functions specified in one process or multiple processes in the flowchart and/or one block or multiple blocks in the block diagram.
[0195] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing equipment to work in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture including the instruction device. , The instruction device realizes the functions specified in one process or multiple processes in the flowchart and/or one block or multiple blocks in the block diagram.
[0196] These computer program instructions can also be loaded on a computer or other programmable data processing equipment, so that a series of operation steps are executed on the computer or other programmable equipment to produce computer-implemented processing, so that the computer or other programmable equipment The instructions executed above provide steps for implementing functions specified in a flow or multiple flows in the flowchart and/or a block or multiple blocks in the block diagram.
[0197] The above content is only to illustrate the technical ideas of the present invention, and cannot be used to limit the scope of protection of the present invention. Any changes made on the basis of the technical solutions based on the technical ideas proposed by the present invention fall into the rights of the present invention. Within the scope of protection of the request.
5 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5
Every citation, both ways
| Document | Relation | Office | Category | Cited during | Relevant claims |
|---|---|---|---|---|---|
| CN119255321A | Cited by | China | – | Search report | – |
| CN115866574A | Cited by | China | – | Search report | – |
| CN117858105A | Cited by | China | – | Search report | – |
| CN116709238A | Cited by | China | – | Search report | – |
| CN114142898A | Cited by | China | – | Search report | – |
| CN114051204A | Cited by | China | – | Search report | – |
| CN119519887A | Cited by | China | – | Search report | – |
| CN113949474A | Cited by | China | – | Search report | – |
| CN114422060A | Cited by | China | – | Search report | – |
| CN116546510A | Cited by | China | – | Search report | – |
| CN117103282A | Cited by | China | – | Search report | – |
| CN114257298A | Cited by | China | – | Search report | – |
| CN114422056A | Cited by | China | – | Search report | – |
| CN116669073A | Cited by | China | – | Search report | – |
| CN115801157A | Cited by | China | – | Search report | – |
| CN114302495A | Cited by | China | – | Search report | – |
| CN114124266A | Cited by | China | – | Search report | – |
| CN115334519A | Cited by | China | – | Search report | – |
| CN114980132A | Cited by | China | – | Search report | – |
| WO2025020222A1 | Cited by | World Intellectual Property Organization (WIPO) | – | International search | – |
| CN111193536A | Cites | China | A | Search report | 1-10 |
| CN111786713A | Cites | China | A | Search report | 1-10 |
| CN112118556A | Cites | China | A | Search report | 1-9 |
| CN112532300A | Cites | China | A | Search report | 1-9 |
| CN112769464A | Cites | China | A | Search report | 1-10 |
| US2013337822A1 | Cites | United States of America | A | Search report | 1-10 |
| US2021003412A1 | Cites | United States of America | A | Search report | 1-9 |
| SHIYU JIAO ET AL: "Joint Beamforming and Phase Shift Design in Downlink UAV Networks with IRS-Assisted NOMA", 《JOURNAL OF COMMUNICATIONS AND INFORMATION NETWORKS》 | Non-patent | – | – | Search report | – |
| CHENGCHENG FENG ET AL: "Trajectory and Beamforming Vector Optimization for Multi-UAV Multicast Network", 《2019 11TH INTERNATIONAL CONFERENCE ON WIRELESS COMMUNICATIONS AND SIGNAL PROCESSING》 | Non-patent | – | – | Search report | – |
| ZINA MOHAMED ET AL: "Resource Allocation for Energy-Efficient Cellular Communications via Aerial IRS", 《2021 IEEE WIRELESS COMMUNICATIONS AND NETWORKING CONFERENCE》 | Non-patent | – | – | Search report | – |
| LINGHUI GE ET AL: "Joint Beamforming and Trajectory Optimization for Intelligent Reflecting Surfaces-Assisted UAV Communications", 《IEEE ACCESS》 | Non-patent | – | – | Search report | – |
| SHENGJUN WU: "Illegal Radio Station Localization with UAV-Based Q-Learning", 《中国通信》 | Non-patent | – | – | Search report | – |
| 郝立元: "无人机中继通信轨迹和功率优化策略研究", 《电子制作》 | Non-patent | – | – | Search report | – |
2 members in 1 office
Members2
| Document | Office | Kind | |
|---|---|---|---|
| CN113364495AThis record | China | A | |
| CN113364495B | China | B |
4 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Termination of patent right due to non-payment of annual feeCF01 | CF01 | |
| Patent grantGrantedGR01 | GR01 | |
| Entry into force of request for substantive examinationSE01 | SE01 | |
| PublicationPB01 | PB01 |
Numbers
- Publication
- 113364495
- Application
- 105730246
Titles2
- Chinese
- 一种多无人机轨迹和智能反射面相移联合优化方法及系统
- English
- Method and system for joint optimization of multi-UAV trajectory and intelligent reflecting surface phase shift
Classification
- CPC, 4
- H04B7/01
- G06F18/23213
- G06F18/214
- Y02D30/70
- IPC, 2
- H04B7 01
- G06K9 62