Efficient AD for the simulation of gradients
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 25/100
- Tipo di issue
- Funzionalità
- Chiarezza
- Da chiarire
- Stato di attività
- Ferma
- Stack tecnologico
- julia
- Ambito
- machine-learning, performance
Direzione di ricerca
Inizia leggendo chainrules.jl e il design dello spazio degli input con output multipli collegato nell’issue. Confronta il percorso autodiff attuale con i casi proposti di kernel isotropi e stazionari, quindi definisci un approccio circoscritto alle derivate del primo ordine e i relativi criteri di validazione; l’issue non indica test o file specifici oltre a chainrules.jl.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Introduction: How would you simulate gradients
If you want to simulate the gradient of a random function $Z$, it turns out that you simply need to take derivatives of the covariance funcion, as (w.l.o.g. $Z$ is centered)
\text{Cov}(\nabla Z(x), Z(y)) = \mathbb{E}[\nabla Z(x) Z(y)] = \nabla_x \mathbb{E}[Z(x)Z(y)] = \nabla_x C(x,y)
And similarly
\text{Cov}(\nabla Z(x), \nabla Z(y)) = \nabla_x\nabla_y C(x,y)
For stationary covariance functions $C(x,y) = C(x-y)$ this simplifies to
\begin{aligned}
\nabla_x C(x,y) &= \nabla C(x-y)\\
\nabla_x\nabla_y C(x,y) &= -\nabla^2 C(x-y)
\end{aligned}
So if we consider the multivariate random function $T_1(x) = (Z(x), \nabla Z(x))$, then its covariance kernel is given by
\begin{aligned}
C_{T_1}(x,y)
&= \begin{pmatrix}
C(x,y) & \nabla_y C(x,y)^T\\
\nabla_x C(x,y) & \nabla_x \nabla_y C(x,y)
\end{pmatrix}\\
&\overset{\mathllap{\text{stationary}}}=
\begin{pmatrix}
C(x-y) & -\nabla C(x-y)^T\\
\nabla C(x-y) & -\nabla^2C(x-y)
\end{pmatrix}\\
&\overset{\mathllap{\text{isotrope}}}=
\begin{pmatrix}
C(d) & -f'\bigl(\frac{\|d\|^2}{2}\bigr)d^T\\
f'\bigl(\frac{\|d\|^2}{2}\bigr) d & -\Bigl[f''\bigl(\frac{\|d\|^2}2\bigr)dd^T + f'\bigl(\frac{\|d\|^2}2\bigr) \mathbb{I}\Bigr]
\end{pmatrix}
\end{aligned}
where we use $d=x-y$ and $C(d) = f\bigl(\frac{|d|^2}2\bigr)$ in the last equation.
Performance Considerations
In principle you could just directly apply Autodiff (AD) to any kernel $C$ to obtain $C_{T_1}$, but since I heard that the expense of autodiff scales with the number of input arguments, this would be really wasteful when we only really need to differentiate a one dimensional function in the isotropic case.
Unfortunately the way that KernelFunctions implement length scales results in general kernel functions, so I am not completely sure how to tell the compiler, that "these derivatives are much simpler than they look".
One possibility might be, to add the Abstract types IsotropicKernel and StationaryKernel and carry these types over when transformations do not violate them. Scaling would not, more general affine transformations would violate isotropy but not stationarity, etc. This could probably be done with type parameters.
But even once that is implemented, how do you tell autodiff what to differentiate? I have seen the file chainrules.jl in this repository, so I thought I would ask if someone already knows how to implement this.
Kernels for multiple outputs considerations
Since you implemented kernels for multiple outputs as an extension of the input space, the reuse of the derivative $f'(\frac{|d|}{2})$ also becomes more complicated.
Maybe this is all premature optimization, as the evaluation of the kernel is complexity wise in the shadow of the cholesky decomposition.
Extension: Simulate $n$-order derivatives
In principle you could similarly simulate $T_n(x) = (Z(x), Z'(x), \dots, Z^{(n)}(x))$. But AD is already underdeveloped for hessians, so I don't know how to get this to work. It might be possible to write custom logic for certain isotropic random functions, like with the squared exponential one.
What do you think? I handcrafted something for first order derivatives in a personal project, but for KernelFunctions.jl a more general approach is probably needed.
- Lingua principale
- Julia
- Stelle
- 276
- Fork
- 41
- Merge medio
- 1g 13h
- PR unite (30g)
- 1
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di JuliaGaussianProcesses/KernelFunctions.jl
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
JuliaGaussianProcesses/KernelFunctions.jl#609 · 1 commento ·
-
Future plans?Aperta
Difficoltà 5/5 Più di una settimana Idoneità per principianti 25/100
JuliaGaussianProcesses/KernelFunctions.jl#585 · 8 commenti · 3 reazioni ·
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 30/100
-
Difficoltà 3/5 1-2 giorni Idoneità per principianti 48/100
JuliaGaussianProcesses/KernelFunctions.jl#571 · 1 reazione ·
-
new kernel
Difficoltà 5/5 Più di una settimana Idoneità per principianti 25/100
Tutte le issue di JuliaGaussianProcesses/KernelFunctions.jl
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
Sienna-Platform/PowerSystemCaseBuilder.jl#239 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
JuliaPluto/Malt.jl#114 ·
-
enhancement
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
QuantumSavory/QuantumSavory.jl#592 ·
I maintainer di solito rispondono entro 1 giorno
-
bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
JuliaQUBO/QUBOTools.jl#136 ·