[Question] windows([1,k,k]) vs. windows([k,k]) performance gap
Dieses Issue hat noch niemand übernommen.
Bewertung
- Schwierigkeit
- 4/5
- Geschätzter Aufwand
- 3-5 Tage
- Anfängerfreundlichkeit
- 35/100
- Issue-Typ
- Bug
- Klarheit
- Größtenteils klar
- Aktivitätsstatus
- Veraltet
- Tech-Stack
- rust
- Bereich
- performance
Rechercherichtung
Start with the windows() call and the two Rust code paths shown in the issue. Build the promised minimal benchmark for [1, k, k] versus [k, k] in release mode, ensuring the compiler cannot optimize away the work. Compare the iterator, into_owned, into_shape, and row assignment costs; done means identifying the source of the gap and either fixing it or documenting the cause.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Beschreibung
For my convolution preprocessing I'm using the windows() method.
Lately I've been doing some refactoring to get my api closer the the one from pytorch.
While doing so I had to change my mnist training image-ndarray from [60_000, 28,28] to [60_000, 1, 28, 28].
Therefore I had to use the windows of the shape [1, k, k] instead of [k, k]. (We iterate over the 60k examples)
My runtime on a test-set increased by roughly 20% from 3.3s to 4s.
Another example (cifar10) had been using the 3d part before and was unaffected by the refactoring,
so I suspect the windows() method to cause the difference.
I am using Rust 1.48 in release-mode and ndarray 13.1.
I will later create and add a mwe repository later, I still have to figure out a benchmark which isn't optimized away by the compiler.
However, here is already the relevant code:
if input.ndim() == 2 {
// this path was used for mnist before refactoring
let x_2d: Array2<f32> = input.into_dimensionality::<Ix2>().unwrap();
let windows = x_2d.windows([k, k]);
for window in windows {
let unrolled: Array1<f32> = window
.into_owned()
.into_shape(k * k * filter_depth)
.unwrap();
xx.row_mut(row_num).assign(&unrolled);
row_num += 1;
}
} else {
// this path is used after refactoring, filter_depth == 1 in this case.
let x_3d: Array3<f32> = input.into_dimensionality::<Ix3>().unwrap();
if forward {
let windows = x_3d.windows([filter_depth, k, k]);
for window in windows {
let unrolled: Array1<f32> = window
.into_owned()
.into_shape(k * k * filter_depth)
.unwrap();
xx.row_mut(row_num).assign(&unrolled);
row_num += 1;
}
}
Given that the first dim of the window and the first dim of the input are both equal and one, I would expect the performance to stay almost equal, since we iterate over the same amount of elements. I guess that having an extra (zero) dimension shouldn't impact the iterator performance so much, since we at the most have a single extra layer of indirection?
I'm generally assuming that the performance of into_shape stays the same, since the elements should have the same memory layout (except maybe of an additional layer of indirection for the 3d case).
In my own use case I could probably reshape each (1,28,28) array to (28,28) and add the extra dimension back somewhere down the line. However, it might (or might not) be interesting to others, so if you are interested @xd009642 @bluss I could try to fix that.
Checking each input in windows() for such a case, downshaping the array if needed, and add an extra dimension back to each iterator-step output would probably be inefficient, since it would require a reshape+unwrap call for each single window.
So I guess the right way would be to find out which ndarray part takes longer and fix the issue there.
What are you expecting as a reason, is creating a higher-dim arrayview so complex that it takes 20% longer?
Or should I start by looking somewhere else?
- Vorherrschende Sprache
- Rust
- Sterne
- 4.3k
- Forks
- 391
- PR-Merge-Kennzahlen
- Keine gemergten PRs in 30 T.
Beitragsleitfaden
Für dieses Repository ist kein Beitragsleitfaden indexiert
Erste Schritte
- Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
- Forken Sie das Repository und arbeiten Sie in einem Branch.
- Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.
Mehr aus rust-ndarray/ndarray
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 72/100
rust-ndarray/ndarray#1612 · 1 Kommentar ·
-
Schwierigkeit 4/5 3-5 Tage Anfängerfreundlichkeit 48/100
rust-ndarray/ndarray#1617 · 1 Kommentar ·
-
Stack overflow in `triu` Offenbug good first issue
Schwierigkeit 3/5 1-2 Tage Anfängerfreundlichkeit 68/100
rust-ndarray/ndarray#1615 · 1 Kommentar ·
-
Schwierigkeit 4/5 3-5 Tage Anfängerfreundlichkeit 48/100
rust-ndarray/ndarray#1610 ·
-
Schwierigkeit 3/5 1-2 Tage Anfängerfreundlichkeit 72/100
rust-ndarray/ndarray#1609 ·
Alle Issues in rust-ndarray/ndarray
Ähnliche Issues
-
bug
Schwierigkeit 1/5 Unter einer Stunde Anfängerfreundlichkeit 85/100
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 75/100
yantrikos/yantrik-os#255 ·
-
bug CLI custom-model
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 75/100
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 70/100
raphamorim/rio#1956 ·
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 75/100
rust-bitcoin/rust-bitcoin#6930 · 1 Kommentar ·