Sentence Similarity
sentence-transformers
Safetensors
feature-extraction
Generated from Trainer
dataset_size:100000
loss:MultipleNegativesRankingLoss
Eval Results (legacy)
Instructions to use duckduckpuck/sir-sbert-e5-large-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use duckduckpuck/sir-sbert-e5-large-v1 with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("duckduckpuck/sir-sbert-e5-large-v1") sentences = [ "Query: The Gaumont Film Company was founded before a studio that was established in Denmark by what Danish filmmaker?\n\nContext:\n\n1. It is the first and oldest film company in the world, founded before other studios such as Pathé (founded in 1896), Titanus (1904), Nordisk Film (1906), Universal and Paramount Pictures (both founded in 1912).", "Nordisk Film (or Nordisk Film Distribution, USA affiliate: Great Northern Film Company), established in Denmark in 1906 by Danish filmmaker Ole Olsen and also the oldest continuously active film studio in the world. It is the third oldest studio in the world behind the Gaumont Film Company and Pathé. Olsen started his company in the Copenhagen suburb of Valby under the name \"Ole Olsen's Film Factory\" but soon changed it to the Nordisk Film Kompagni. In 1908, Olsen opened an affiliate branch in New York, the Great Northern Film Company, to handle distribution of his films to the American market. As Nordisk Film, it became a publicly traded company in 1911.", "The West Lodge, also known as the West Gate Lodge, to Cardiff Castle is a Grade II* listed building, currently used as a tea room, in the centre of Cardiff, Wales. It is approximately 100 m west of the Castle, with the Animal Wall running in-between.", "Julmust (Swedish: \"jul\" \"Yule\" and \"must \" \"not yet fermented juice of fruit or berries\", though there is no such juice in \"julmust\") is a soft drink that is mainly consumed in Sweden around Christmas. During the other part of the year it is usually quite difficult to find in stores, but sometimes it is sold at other times of the year under the name \"must\". At Easter the name is påskmust (from \"påsk \", \"Easter\" / \"Paschal\" [\"q.v.\"]). The content is the same regardless of the marketing name, although the length of time it is stored before bottling differs; however, the beverage is more closely associated with Christmas, somewhat less with Easter and traditionally not at all with the summer. 45 million litres of \"julmust\" are consumed during December, which is around 50% of the total soft drink volume in December and 75% of the total yearly must sales." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
SentenceTransformer based on embaas/sentence-transformers-e5-large-v2
This is a sentence-transformers model finetuned from embaas/sentence-transformers-e5-large-v2. It maps sentences & paragraphs to a 1024-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.
Model Details
Model Description
- Model Type: Sentence Transformer
- Base model: embaas/sentence-transformers-e5-large-v2
- Maximum Sequence Length: 512 tokens
- Output Dimensionality: 1024 dimensions
- Similarity Function: Cosine Similarity
Model Sources
- Documentation: Sentence Transformers Documentation
- Repository: Sentence Transformers on GitHub
- Hugging Face: Sentence Transformers on Hugging Face
Full Model Architecture
SentenceTransformer(
(0): Transformer({'max_seq_length': 512, 'do_lower_case': False}) with Transformer model: PeftModelForFeatureExtraction
(1): Pooling({'word_embedding_dimension': 1024, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
(2): Normalize()
)
Model tree for duckduckpuck/sir-sbert-e5-large-v1
Base model
embaas/sentence-transformers-e5-large-v2Evaluation results
- Cosine Accuracy@1 on eval splitself-reported0.923
- Cosine Accuracy@3 on eval splitself-reported0.987
- Cosine Accuracy@5 on eval splitself-reported0.991
- Cosine Accuracy@10 on eval splitself-reported0.995
- Cosine Precision@1 on eval splitself-reported0.923
- Cosine Precision@3 on eval splitself-reported0.329
- Cosine Precision@5 on eval splitself-reported0.198
- Cosine Precision@10 on eval splitself-reported0.099