Poor slicing performance compared to NumPy

Aberta
#84 10 comentários 0 reações 0 responsáveis Ver no GitHub

Ninguém assumiu esta issue ainda.

Avaliação

Dificuldade
4/5
Tempo estimado
3-5 dias
Facilidade para iniciantes
25/100
Tipo de issue
Bug
Clareza
Precisa de esclarecimento
Status de atividade
Estagnada
Stack de tecnologia
python
Domínio
performance

Direção de pesquisa

Comece executando o benchmark Python fornecido com os backends de CPU e GPU, concentrando-se em af_B[:, i], af_B[i, :] e af.matmul. Como nenhum arquivo-fonte ou teste é mencionado na issue, rastreie esses pontos de entrada pelos bindings de Python e compare-os com unsliced matmul. O trabalho estará concluído quando você reproduzir a regressão, identificar sua causa e adicionar um teste de regressão ou benchmark que demonstre a melhoria.

Escrita pelo modelo de indexação a partir do texto da issue.

Descrição

reported by @floopcz on over here: https://github.com/arrayfire/arrayfire/issues/1428

ArrayFire slicing seems to suffer from a performance issue. Consider the following python code, that:

  • calculates the dot product of two matrices, first using NumPy, than ArrayFire
  • calculates each column/row of the dot product separately by slicing a single column/row from one of the matrices
#!/usr/bin/env python3
from time import time
import arrayfire as af
import numpy as np

af.set_backend('cpu')
af.info()

iters = 1000
n = 512

af_A = af.randu(n, n)
af_B = af.randu(n, n)

np_A = np.random.rand(n, n).astype(np.float32)
np_B = np.random.rand(n, n).astype(np.float32)

start = time()
for t in range(iters):
    np_C = np.dot(np_A, np_B)
print('numpy - dot: {}'.format(time() - start))

af.sync()
start = time()
for t in range(iters):
    af_C = af.matmul(af_A, af_B)
af.sync()
print('arrayfire - matmul: {}'.format(time() - start))

start = time()
for t in range(iters):
    for i in range(np_B.shape[1]):
        np_C = np.dot(np_A, np_B[:, i])
print('numpy - sliced dot - column major: {}'.format(time() - start))

af.sync()
start = time()
for t in range(iters):
    for i in range(af_B.shape[1]):
        af_C = af.matmul(af_A, af_B[:, i])
af.sync()
print('arrayfire - sliced matmul - column major: {}'.format(time() - start))

start = time()
for t in range(iters):
    for i in range(np_B.shape[0]):
        np_C = np.dot(np_B[i, :], np_A)
print('numpy - sliced dot - row major: {}'.format(time() - start))

af.sync()
start = time()
for t in range(iters):
    for i in range(af_B.shape[0]):
        af_C = af.matmul(af_B[i, :], af_A)
af.sync()
print('arrayfire - sliced matmul - row major: {}'.format(time() - start))

The results are following:

ArrayFire v3.3.2 (CPU, 64-bit Linux, build f65dd97)
[0] Unknown: Unknown, 15880 MB, Max threads(1) 
numpy - dot: 1.3848536014556885
arrayfire - matmul: 1.325775146484375
numpy - sliced dot - column major: 7.156768798828125
arrayfire - sliced matmul - column major: 38.87605834007263
numpy - sliced dot - row major: 7.6784679889678955
arrayfire - sliced matmul - row major: 41.27544379234314

The results suggest that with slicing, arrayfire performance is significantly degraded compared to NumPy. I have achieved similarly distributed results also with the GPU backend. Both numpy and arrayfire are linked against Intel MKL.

Am I doing something "illegal" or is it an inefficiency of the library? Thanks.

Linguagem predominante
Python
Estrelas
422
Forks
63
Métricas de merge de PRs
Nenhum PR com merge em 30d

Guia de contribuição

Nenhum guia de contribuição indexado para este repositório

Primeiros passos

  1. Leia a issue inteira e depois o guia de contribuição do projeto.
  2. Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
  3. Faça um fork do repositório e trabalhe em uma branch.
  4. Abra um pull request que referencie o número da issue.

Mais de arrayfire/arrayfire-python

Todas as issues de arrayfire/arrayfire-python

Issues semelhantes

Mais issues de Python

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.