Supporting non-vision models
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 25/100
- Tipo de issue
- Nueva funcionalidad
- Claridad
- Necesita aclaración
- Estado de actividad
- Estancado
- Área
- machine-learning
Línea de trabajo
Comienza siguiendo get_previous_layer hasta la llamada a re.sub que recibe None y, a continuación, compara sus suposiciones sobre Conv2d/BatchNorm2d con las cadenas reportadas de Conv1d, Linear, LayerNorm, GELU, transpose, dropout y add. Se considerará terminado cuando se documente qué predecesor debe seleccionarse para modelos que no sean de visión y se defina el comportamiento cuando no se encuentre ninguna capa compatible.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Hi,
could you explain how to use this for non-vision models?
Trying to apply this to an audio model I am getting an error because get_previous_layer never finds a Conv2d or a BatchNorm2d and hence, after recursing through all the modules, just returns None, which then fails at re.sub.
Could you explain the logic behind "going back to the previous layer of exactly this type"?
the immediate earlier layers are: (add, causing the search) -> [transpose] -> gelu -> transpose -> layer_norm -> transpose -> conv1d -> ...
which layer would you expect to find here? Are you looking for the last layer that learns anything (which would be layer norm) or with actual learnable parameters (then it's conv1d)
there is a second case where this happens where the chain is: (add, causing the search) -> [dropout] -> linear -> layer_norm -> transpose -> gelu -> transpose -> layer_norm -> transpose -> conv1d -> ...
again, please help me out which layer should be found. I would guess the linear layer?
on a side note: instead of deep recursion, checking the whole list each time, building a tree or doubly-linked-list-like structure seems more appropriate (and easier to debug)
- Lenguaje dominante
- Python
- Estrellas
- 36
- Forks
- 3
- Métricas de merge de PR
- Sin PR fusionados en 30 d
Preparar el entorno
Este proyecto no incluye contenedor de desarrollo, Dockerfile ni guía de contribución, así que la configuración corre por tu cuenta: empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de EIDOSLAB/simplify
-
Make into Conda packageAbierto
Dificultad 3/5 1-2 días Aptitud para principiantes 25/100
-
Not work for Swin_T modelAbierto
Dificultad 4/5 3-5 días Aptitud para principiantes 25/100
-
Accuracy drop after simplifyAbierto
Dificultad 4/5 3-5 días Aptitud para principiantes 25/100
-
Big accuracy drop after simplifyQuizá libre de nuevo @AndreaBrg la tomó hace 1330 días y no hay ningún pull request abierto. Abiertobug
Todos los issues de EIDOSLAB/simplify
Issues similares
-
needs-human needs-triage
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
gke-labs/kube-agents#2400 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Device Details tables: FS/SF columns contradict each other (nfet_01v8 Vt row, pfet_01v8 Idsat row)Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
google/skywater-pdk#450 ·
-
Drained trajectory arrays are overwritten when the sequence buffer is reusedPosiblemente ocupada @sylvesterkaczmarek la tomó hoy. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
google-deepmind/bsuite#56 ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
LearningCircuit/local-deep-research#7206 ·
Los mantenedores suelen responder en 1 día
-
[TASK] Document technology stackAbierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
chingu-voyages/V62-tier3-team-33#285 ·
Los mantenedores suelen responder en 1 día