felflare
/

bert-restore-punctuation

@@ -12,13 +12,13 @@ The model predicts the punctuation and upper-casing of plain, lower-cased text.
 This model is intended for direct use as a punctuation restoration model for the general English language. Alternatively, you can use this for further fine-tuning on domain-specific texts for punctuation restoration tasks.
-Model restores the following punctuations -- [` ! ? . , - : ; '`]
-Model also restores upper-casing of words.
 -----------------------------------------------
 ## 🚋 Usage
-Below is a quick way to get up and running with the model.
 1. First, install the package.
 ```bash
 pip install rpunct
@@ -28,24 +28,31 @@ pip install rpunct
 from rpunct import RestorePuncts
 # The default language is 'english'
 rpunct = RestorePuncts()
-rpunct.punctuate("""in 2018 cornell researchers built a high-powered detector that in combination with an algorithm-driven process called ptychography set a world record by tripling the resolution of a state-of-the-art electron microscope as successful as it was that approach had a weakness it only worked with ultrathin samples that were a few atoms thick anything thicker would cause the electrons to scatter in ways that could not be disentangled now a team again led by david muller the samuel b eckert professor of engineering has bested its own record by a factor of two with an electron microscope pixel array detector empad that incorporates even more sophisticated 3d reconstruction algorithms the resolution is so fine-tuned the only blurring that remains is the thermal jiggling of the atoms themselves""")
 # Outputs the following:
-# In 2018, Cornell researchers built a high-powered detector that, in combination with an algorithm-driven process called Ptychography, set a world record by tripling the resolution of a state-of-the-art electron microscope. As successful as it was, that approach had a weakness. It only worked with ultrathin samples that were a few atoms thick. Anything thicker would cause the electrons to scatter in ways that could not be disentangled. Now, a team again led by David Muller, the Samuel B. Eckert Professor of Engineering, has bested its own record by a factor of two with an Electron microscope pixel array detector empad that incorporates even more sophisticated 3d reconstruction algorithms. The resolution is so fine-tuned the only blurring that remains is the thermal jiggling of the atoms themselves.
 ```
-`This model works on arbitrarily large text in English language and uses GPU if available.`
 -----------------------------------------------
 ## 📡 Training data
 Here is the number of product reviews we used for finetuning the model:
-| Language | Number of reviews |
 | -------- | ----------------- |
 | English  | 560,000           |
-We found the best convergence around `3 epochs`, which is what presented here and available via a download.
 -----------------------------------------------
 ## 🎯 Accuracy
@@ -76,7 +83,6 @@ Below is a breakdown of the performance of the model by each label:
 |     **Upper**    |   0.84       | 0.82    |  0.83     | 5442
 -----------------------------------------------
 ## ☕ Contact
 Contact [Daulet Nurmanbetov]([email protected]) for questions, feedback and/or requests for similar models.

 This model is intended for direct use as a punctuation restoration model for the general English language. Alternatively, you can use this for further fine-tuning on domain-specific texts for punctuation restoration tasks.
+Model restores the following punctuations -- **[! ? . , - : ; ' ]**
+The model also restores the upper-casing of words.
 -----------------------------------------------
 ## 🚋 Usage
+**Below is a quick way to get up and running with the model.**
 1. First, install the package.
 ```bash
 pip install rpunct
 from rpunct import RestorePuncts
 # The default language is 'english'
 rpunct = RestorePuncts()
+rpunct.punctuate("""in 2018 cornell researchers built a high-powered detector that in combination with an algorithm-driven process called ptychography set a world record
+by tripling the resolution of a state-of-the-art electron microscope as successful as it was that approach had a weakness it only worked with ultrathin samples that were
+a few atoms thick anything thicker would cause the electrons to scatter in ways that could not be disentangled now a team again led by david muller the samuel b eckert
+professor of engineering has bested its own record by a factor of two with an electron microscope pixel array detector empad that incorporates even more sophisticated
+3d reconstruction algorithms the resolution is so fine-tuned the only blurring that remains is the thermal jiggling of the atoms themselves""")
 # Outputs the following:
+# In 2018, Cornell researchers built a high-powered detector that, in combination with an algorithm-driven process called Ptychography, set a world record by tripling the
+# resolution of a state-of-the-art electron microscope. As successful as it was, that approach had a weakness. It only worked with ultrathin samples that were a few atoms
+# thick. Anything thicker would cause the electrons to scatter in ways that could not be disentangled. Now, a team again led by David Muller, the Samuel B.
+# Eckert Professor of Engineering, has bested its own record by a factor of two with an Electron microscope pixel array detector empad that incorporates even more
+# sophisticated 3d reconstruction algorithms. The resolution is so fine-tuned the only blurring that remains is the thermal jiggling of the atoms themselves.
 ```
+**This model works on arbitrarily large text in English language and uses GPU if available.**
 -----------------------------------------------
 ## 📡 Training data
 Here is the number of product reviews we used for finetuning the model:
+| Language | Number of text samples|
 | -------- | ----------------- |
 | English  | 560,000           |
+We found the best convergence around _**3 epochs**_, which is what presented here and available via a download.
 -----------------------------------------------
 ## 🎯 Accuracy
 |     **Upper**    |   0.84       | 0.82    |  0.83     | 5442
 -----------------------------------------------
 ## ☕ Contact
 Contact [Daulet Nurmanbetov]([email protected]) for questions, feedback and/or requests for similar models.